{"thread":{"id":"65563","subject":"[PATCH 0/6] Handle cloning of objects larger than 4GB on Windows","startedAt":"2026-04-28T16:26:24Z","lastAt":"2026-05-11T10:02:20Z","messageCount":60,"participants":["Johannes Schindelin via GitGitGadget","Derrick Stolee","Torsten Bögershausen","Jeff King","Johannes Schindelin","Junio C Hamano","Patrick Steinhardt"],"isPatch":true,"patchVersion":1,"patchTotal":6},"messages":[{"id":"542435","messageId":"pull.2102.git.1777393580.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":null,"subject":"[PATCH 0/6] Handle cloning of objects larger than 4GB on Windows","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:14Z","receivedAt":"2026-04-28T16:26:24Z","isPatch":true,"body":"On Windows, unsigned long is 32-bit even on 64-bit systems. This causes\nmultiple problems when Git handles objects larger than 4GB. This patch\nseries is a very targeted fix for a very early part of the problem: it\naddresses the most fundamental truncation points that prevent a >4GB object\nfrom surviving a clone at all.\n\nSpecifically, this fixes:\n\n * zlib's uLong wrapping and triggering BUG() assertions in the git_zstream\n   wrapper\n * Object sizes being truncated in pack streaming, delta headers, and\n   index-pack/unpack-objects\n * pack-objects re-encoding reused pack entries with a truncated size,\n   producing corrupt packs on the wire\n\nMany other code paths still use unsigned long for object sizes (e.g.,\ncat-file -s, object_info.sizep, the delta machinery) and will need their own\nconversions. This series does not attempt to fix those.\n\nBased on work by @LordKiRon in git-for-windows/git#6076.\n\nThe last two commits add a test helper that synthesizes a pack with a >4GB\nblob and regression tests that clone it via both the unpack-objects and\nindex-pack code paths using file:// transport.\n\nJohannes Schindelin (6):\n  index-pack, unpack-objects: use size_t for object size\n  git-zlib: handle data streams larger than 4GB\n  odb, packfile: use size_t for streaming object sizes\n  delta, packfile: use size_t for delta header sizes\n  test-tool: add a helper to synthesize large packfiles\n  t5608: add regression test for >4GB object clone\n\n Makefile                     |   1 +\n builtin/index-pack.c         |   9 +-\n builtin/pack-objects.c       |  23 +++-\n builtin/unpack-objects.c     |   5 +-\n compat/zlib-compat.h         |   2 +\n delta.h                      |  14 +-\n git-zlib.c                   |  25 ++--\n git-zlib.h                   |   4 +-\n object-file.c                |  12 +-\n odb/streaming.c              |  13 +-\n odb/streaming.h              |   2 +-\n oss-fuzz/fuzz-pack-headers.c |   2 +-\n pack-bitmap.c                |   2 +-\n pack-check.c                 |   6 +-\n packfile.c                   |  57 +++++---\n packfile.h                   |   4 +-\n t/helper/meson.build         |   1 +\n t/helper/test-synthesize.c   | 250 +++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c         |   1 +\n t/helper/test-tool.h         |   1 +\n t/t5608-clone-2gb.sh         |  37 ++++++\n 21 files changed, 418 insertions(+), 53 deletions(-)\n create mode 100644 t/helper/test-synthesize.c\n\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2102%2Fdscho%2Ffix-large-clones-on-windows-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2102/dscho/fix-large-clones-on-windows-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/2102\n-- \ngitgitgadget\n"},{"id":"542436","messageId":"dc660106ea8511e6adc44d2b70e9a4ae8b18090e.1777393580.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH 1/6] index-pack, unpack-objects: use size_t for object size","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:15Z","receivedAt":"2026-04-28T16:26:25Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen unpacking objects from a packfile, the object size is decoded\nfrom a variable-length encoding. On platforms where unsigned long is\n32-bit (such as Windows, even in 64-bit builds), the shift operation\noverflows when decoding sizes larger than 4GB. The result is a\ntruncated size value, causing the unpacked object to be corrupted or\nrejected.\n\nFix this by changing the size variable to size_t, which is 64-bit on\n64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\nspace.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n builtin/index-pack.c     | 9 +++++----\n builtin/unpack-objects.c | 5 +++--\n 2 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..cc660582e9 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -37,7 +37,7 @@ static const char index_pack_usage[] =\n \n struct object_entry {\n \tstruct pack_idx_entry idx;\n-\tunsigned long size;\n+\tsize_t size;\n \tunsigned char hdr_size;\n \tsigned char type;\n \tsigned char real_type;\n@@ -469,7 +469,7 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n-static void *unpack_entry_data(off_t offset, unsigned long size,\n+static void *unpack_entry_data(off_t offset, size_t size,\n \t\t\t       enum object_type type, struct object_id *oid)\n {\n \tstatic char fixed_buf[8192];\n@@ -524,7 +524,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \t\t\t      struct object_id *oid)\n {\n \tunsigned char *p;\n-\tunsigned long size, c;\n+\tsize_t size;\n+\tunsigned long c;\n \toff_t base_offset;\n \tunsigned shift;\n \tvoid *data;\n@@ -542,7 +543,7 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \t\tp = fill(1);\n \t\tc = *p;\n \t\tuse(1);\n-\t\tsize += (c & 0x7f) << shift;\n+\t\tsize += ((size_t)c & 0x7f) << shift;\n \t\tshift += 7;\n \t}\n \tobj->size = size;\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..59a36c2481 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -533,7 +533,8 @@ static void unpack_one(unsigned nr)\n {\n \tunsigned shift;\n \tunsigned char *pack;\n-\tunsigned long size, c;\n+\tsize_t size;\n+\tunsigned long c;\n \tenum object_type type;\n \n \tobj_list[nr].offset = consumed_bytes;\n@@ -548,7 +549,7 @@ static void unpack_one(unsigned nr)\n \t\tpack = fill(1);\n \t\tc = *pack;\n \t\tuse(1);\n-\t\tsize += (c & 0x7f) << shift;\n+\t\tsize += ((size_t)c & 0x7f) << shift;\n \t\tshift += 7;\n \t}\n \n-- \ngitgitgadget\n\n"},{"id":"542437","messageId":"92f4327b1fe09126dd6421b071a9071ce5530371.1777393580.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH 2/6] git-zlib: handle data streams larger than 4GB","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:16Z","receivedAt":"2026-04-28T16:26:26Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOn Windows, zlib's `uLong` type is 32-bit even on 64-bit systems. When\nprocessing data streams larger than 4GB, the `total_in` and `total_out`\nfields in zlib's `z_stream` structure wrap around, which caused the\nsanity checks in `zlib_post_call()` to trigger `BUG()` assertions.\n\nThe git_zstream wrapper now tracks its own 64-bit totals rather than\ncopying them from zlib. The sanity checks compare only the low bits,\nusing `maximum_unsigned_value_of_type(uLong)` to mask appropriately for\nthe platform's `uLong` size.\n\nThis is based on work by LordKiRon in git-for-windows#6076.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n git-zlib.c    | 25 +++++++++++++++++--------\n git-zlib.h    |  4 ++--\n object-file.c |  2 +-\n 3 files changed, 20 insertions(+), 11 deletions(-)\n\ndiff --git a/git-zlib.c b/git-zlib.c\nindex df9604910e..b91cb323ae 100644\n--- a/git-zlib.c\n+++ b/git-zlib.c\n@@ -30,6 +30,9 @@ static const char *zerr_to_string(int status)\n  */\n /* #define ZLIB_BUF_MAX ((uInt)-1) */\n #define ZLIB_BUF_MAX ((uInt) 1024 * 1024 * 1024) /* 1GB */\n+\n+/* uLong is 32-bit on Windows, even on 64-bit systems */\n+#define ULONG_MAX_VALUE maximum_unsigned_value_of_type(uLong)\n static inline uInt zlib_buf_cap(unsigned long len)\n {\n \treturn (ZLIB_BUF_MAX < len) ? ZLIB_BUF_MAX : len;\n@@ -39,31 +42,37 @@ static void zlib_pre_call(git_zstream *s)\n {\n \ts->z.next_in = s->next_in;\n \ts->z.next_out = s->next_out;\n-\ts->z.total_in = s->total_in;\n-\ts->z.total_out = s->total_out;\n+\ts->z.total_in = (uLong)(s->total_in & ULONG_MAX_VALUE);\n+\ts->z.total_out = (uLong)(s->total_out & ULONG_MAX_VALUE);\n \ts->z.avail_in = zlib_buf_cap(s->avail_in);\n \ts->z.avail_out = zlib_buf_cap(s->avail_out);\n }\n \n static void zlib_post_call(git_zstream *s, int status)\n {\n-\tunsigned long bytes_consumed;\n-\tunsigned long bytes_produced;\n+\tsize_t bytes_consumed;\n+\tsize_t bytes_produced;\n \n \tbytes_consumed = s->z.next_in - s->next_in;\n \tbytes_produced = s->z.next_out - s->next_out;\n-\tif (s->z.total_out != s->total_out + bytes_produced)\n+\t/*\n+\t * zlib's total_out/total_in are uLong which may wrap for >4GB.\n+\t * We track our own totals and verify only the low bits match.\n+\t */\n+\tif ((s->z.total_out & ULONG_MAX_VALUE) !=\n+\t    ((s->total_out + bytes_produced) & ULONG_MAX_VALUE))\n \t\tBUG(\"total_out mismatch\");\n \t/*\n \t * zlib does not update total_in when it returns Z_NEED_DICT,\n \t * causing a mismatch here. Skip the sanity check in that case.\n \t */\n \tif (status != Z_NEED_DICT &&\n-\t    s->z.total_in != s->total_in + bytes_consumed)\n+\t    (s->z.total_in & ULONG_MAX_VALUE) !=\n+\t    ((s->total_in + bytes_consumed) & ULONG_MAX_VALUE))\n \t\tBUG(\"total_in mismatch\");\n \n-\ts->total_out = s->z.total_out;\n-\ts->total_in = s->z.total_in;\n+\ts->total_out += bytes_produced;\n+\ts->total_in += bytes_consumed;\n \t/* zlib-ng marks `next_in` as `const`, so we have to cast it away. */\n \ts->next_in = (unsigned char *) s->z.next_in;\n \ts->next_out = s->z.next_out;\ndiff --git a/git-zlib.h b/git-zlib.h\nindex 0e66fefa8c..44380e8ad3 100644\n--- a/git-zlib.h\n+++ b/git-zlib.h\n@@ -7,8 +7,8 @@ typedef struct git_zstream {\n \tstruct z_stream_s z;\n \tunsigned long avail_in;\n \tunsigned long avail_out;\n-\tunsigned long total_in;\n-\tunsigned long total_out;\n+\tsize_t total_in;\n+\tsize_t total_out;\n \tunsigned char *next_in;\n \tunsigned char *next_out;\n } git_zstream;\ndiff --git a/object-file.c b/object-file.c\nindex 2acc9522df..086b2b65ff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1118,7 +1118,7 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n \n \tif (stream.total_in != len + hdrlen)\n-\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\tdie(_(\"write stream object %\"PRIuMAX\" != %\"PRIuMAX), (uintmax_t)stream.total_in,\n \t\t    (uintmax_t)len + hdrlen);\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"542438","messageId":"3a539061c5f62c65d46bd0eb774bb1b1239463ff.1777393580.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH 3/6] odb, packfile: use size_t for streaming object sizes","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:17Z","receivedAt":"2026-04-28T16:26:28Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe odb_read_stream structure uses unsigned long for the size field,\nwhich is 32-bit on Windows even in 64-bit builds. When streaming\nobjects larger than 4GB, the size would be truncated to zero or an\nincorrect value, resulting in empty files being written to disk.\n\nChange the size field in odb_read_stream to size_t and introduce\nunpack_object_header_sz() to return sizes via size_t pointer. Since\nobject_info.sizep remains unsigned long for API compatibility, use\ntemporary variables where the types differ, with comments noting the\ntruncation limitation for code paths that still use unsigned long.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n builtin/pack-objects.c       | 23 ++++++++++++++++-------\n object-file.c                | 10 +++++++++-\n odb/streaming.c              | 13 ++++++++++++-\n odb/streaming.h              |  2 +-\n oss-fuzz/fuzz-pack-headers.c |  2 +-\n pack-bitmap.c                |  2 +-\n pack-check.c                 |  6 ++++--\n packfile.c                   | 24 +++++++++++++++---------\n packfile.h                   |  4 ++--\n 9 files changed, 61 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex dd2480a73d..aa4b1cb9b8 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -629,14 +629,21 @@ static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n \tstruct packed_git *p = IN_PACK(entry);\n \tstruct pack_window *w_curs = NULL;\n \tuint32_t pos;\n-\toff_t offset;\n+\toff_t offset, cur;\n \tenum object_type type = oe_type(entry);\n+\tenum object_type in_pack_type;\n \toff_t datalen;\n \tunsigned char header[MAX_PACK_OBJECT_HEADER],\n \t\t      dheader[MAX_PACK_OBJECT_HEADER];\n \tunsigned hdrlen;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n-\tunsigned long entry_size = SIZE(entry);\n+\tsize_t entry_size;\n+\n+\tcur = entry->in_pack_offset;\n+\tin_pack_type = unpack_object_header(p, &w_curs, &cur, &entry_size);\n+\tif (in_pack_type < 0)\n+\t\tdie(_(\"write_reuse_object: unable to parse object header of %s\"),\n+\t\t    oid_to_hex(&entry->idx.oid));\n \n \tif (DELTA(entry))\n \t\ttype = (allow_ofs_delta && DELTA(entry)->idx.offset) ?\n@@ -1087,7 +1094,7 @@ static void write_reused_pack_one(struct packed_git *reuse_packfile,\n {\n \toff_t offset, next, cur;\n \tenum object_type type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \toffset = pack_pos_to_offset(reuse_packfile, pos);\n \tnext = pack_pos_to_offset(reuse_packfile, pos + 1);\n@@ -2243,7 +2250,7 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n \t\toff_t ofs;\n \t\tunsigned char *buf, c;\n \t\tenum object_type type;\n-\t\tunsigned long in_pack_size;\n+\t\tsize_t in_pack_size;\n \n \t\tbuf = use_pack(p, &w_curs, entry->in_pack_offset, &avail);\n \n@@ -2734,16 +2741,18 @@ unsigned long oe_get_size_slow(struct packing_data *pack,\n \tstruct pack_window *w_curs;\n \tunsigned char *buf;\n \tenum object_type type;\n-\tunsigned long used, avail, size;\n+\tunsigned long used, avail;\n+\tsize_t size;\n \n \tif (e->type_ != OBJ_OFS_DELTA && e->type_ != OBJ_REF_DELTA) {\n+\t\tunsigned long sz;\n \t\tpacking_data_lock(&to_pack);\n \t\tif (odb_read_object_info(the_repository->objects,\n-\t\t\t\t\t &e->idx.oid, &size) < 0)\n+\t\t\t\t\t &e->idx.oid, &sz) < 0)\n \t\t\tdie(_(\"unable to get size of %s\"),\n \t\t\t    oid_to_hex(&e->idx.oid));\n \t\tpacking_data_unlock(&to_pack);\n-\t\treturn size;\n+\t\treturn sz;\n \t}\n \n \tp = oe_in_pack(pack, e);\ndiff --git a/object-file.c b/object-file.c\nindex 086b2b65ff..0be2981c7a 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2326,6 +2326,7 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_loose_read_stream *st;\n \tunsigned long mapsize;\n+\tunsigned long size_ul;\n \tvoid *mapped;\n \n \tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n@@ -2349,11 +2350,18 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n \t\tgoto error;\n \t}\n \n-\toi.sizep = &st->base.size;\n+\t/*\n+\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n+\t * st->base.size is size_t (64-bit). Use temporary variable.\n+\t * Note: loose objects >4GB would still truncate here, but such\n+\t * large loose objects are uncommon (they'd normally be packed).\n+\t */\n+\toi.sizep = &size_ul;\n \toi.typep = &st->base.type;\n \n \tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n \t\tgoto error;\n+\tst->base.size = size_ul;\n \n \tst->mapped = mapped;\n \tst->mapsize = mapsize;\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 5927a12954..af2adf5ce7 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -157,15 +157,26 @@ static int open_istream_incore(struct odb_read_stream **out,\n \t\t.base.read = read_istream_incore,\n \t};\n \tstruct odb_incore_read_stream *st;\n+\tunsigned long size_ul;\n \tint ret;\n \n \toi.typep = &stream.base.type;\n-\toi.sizep = &stream.base.size;\n+\t/*\n+\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n+\t * stream.base.size is size_t (64-bit). We use a temporary variable\n+\t * because the types are incompatible. Note: this path still truncates\n+\t * for >4GB objects, but large objects should use pack streaming\n+\t * (packfile_store_read_object_stream) which handles size_t properly.\n+\t * This incore fallback is only used for small objects or when pack\n+\t * streaming is unavailable.\n+\t */\n+\toi.sizep = &size_ul;\n \toi.contentp = (void **)&stream.buf;\n \tret = odb_read_object_info_extended(odb, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n+\tstream.base.size = size_ul;\n \n \tCALLOC_ARRAY(st, 1);\n \t*st = stream;\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex c7861f7e13..517e2ea2d3 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -21,7 +21,7 @@ struct odb_read_stream {\n \todb_read_stream_close_fn close;\n \todb_read_stream_read_fn read;\n \tenum object_type type;\n-\tunsigned long size; /* inflated size of full object */\n+\tsize_t size; /* inflated size of full object */\n };\n \n /*\ndiff --git a/oss-fuzz/fuzz-pack-headers.c b/oss-fuzz/fuzz-pack-headers.c\nindex 150c0f5fa2..ef61ab577c 100644\n--- a/oss-fuzz/fuzz-pack-headers.c\n+++ b/oss-fuzz/fuzz-pack-headers.c\n@@ -6,7 +6,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size);\n int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n {\n \tenum object_type type;\n-\tunsigned long len;\n+\tsize_t len;\n \n \tunpack_object_header_buffer((const unsigned char *)data,\n \t\t\t\t    (unsigned long)size, &type, &len);\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex f6ec18d83a..f9af8a96bd 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -2270,7 +2270,7 @@ static int try_partial_reuse(struct bitmap_index *bitmap_git,\n {\n \toff_t delta_obj_offset;\n \tenum object_type type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \tif (pack_pos >= pack->p->num_objects)\n \t\treturn -1; /* not actually in the pack */\ndiff --git a/pack-check.c b/pack-check.c\nindex 79992bb509..2792f34d25 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -110,7 +110,7 @@ static int verify_packfile(struct repository *r,\n \t\tvoid *data;\n \t\tstruct object_id oid;\n \t\tenum object_type type;\n-\t\tunsigned long size;\n+\t\tsize_t size;\n \t\toff_t curpos;\n \t\tint data_valid;\n \n@@ -143,7 +143,9 @@ static int verify_packfile(struct repository *r,\n \t\t\tdata = NULL;\n \t\t\tdata_valid = 0;\n \t\t} else {\n-\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &size);\n+\t\t\tunsigned long sz;\n+\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &sz);\n+\t\t\tsize = sz;\n \t\t\tdata_valid = 1;\n \t\t}\n \ndiff --git a/packfile.c b/packfile.c\nindex b012d648ad..fdae91dd11 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1133,7 +1133,7 @@ out:\n }\n \n unsigned long unpack_object_header_buffer(const unsigned char *buf,\n-\t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n+\t\tunsigned long len, enum object_type *type, size_t *sizep)\n {\n \tunsigned shift;\n \tsize_t size, c;\n@@ -1144,7 +1144,11 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \tsize = c & 15;\n \tshift = 4;\n \twhile (c & 0x80) {\n-\t\tif (len <= used || (bitsizeof(long) - 7) < shift) {\n+\t\t/*\n+\t\t * Each continuation byte adds 7 bits. Ensure shift won't\n+\t\t * overflow size_t (use size_t not long for 64-bit on Windows).\n+\t\t */\n+\t\tif (len <= used || (bitsizeof(size_t) - 7) < shift) {\n \t\t\terror(\"bad object header\");\n \t\t\tsize = used = 0;\n \t\t\tbreak;\n@@ -1153,7 +1157,7 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \t\tsize = st_add(size, st_left_shift(c & 0x7f, shift));\n \t\tshift += 7;\n \t}\n-\t*sizep = cast_size_t_to_ulong(size);\n+\t*sizep = size;\n \treturn used;\n }\n \n@@ -1215,7 +1219,7 @@ unsigned long get_size_from_delta(struct packed_git *p,\n int unpack_object_header(struct packed_git *p,\n \t\t\t struct pack_window **w_curs,\n \t\t\t off_t *curpos,\n-\t\t\t unsigned long *sizep)\n+\t\t\t size_t *sizep)\n {\n \tunsigned char *base;\n \tunsigned long left;\n@@ -1367,7 +1371,7 @@ static enum object_type packed_to_object_type(struct repository *r,\n \n \twhile (type == OBJ_OFS_DELTA || type == OBJ_REF_DELTA) {\n \t\toff_t base_offset;\n-\t\tunsigned long size;\n+\t\tsize_t size;\n \t\t/* Push the object we're going to leave behind */\n \t\tif (poi_stack_nr >= poi_stack_alloc && poi_stack == small_poi_stack) {\n \t\t\tpoi_stack_alloc = alloc_nr(poi_stack_nr);\n@@ -1586,7 +1590,7 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n \t\t\t\t\t     uint32_t *maybe_index_pos, struct object_info *oi)\n {\n \tstruct pack_window *w_curs = NULL;\n-\tunsigned long size;\n+\tsize_t size;\n \toff_t curpos = obj_offset;\n \tenum object_type type = OBJ_NONE;\n \tuint32_t pack_pos;\n@@ -1778,7 +1782,7 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n \tstruct pack_window *w_curs = NULL;\n \toff_t curpos = obj_offset;\n \tvoid *data = NULL;\n-\tunsigned long size;\n+\tsize_t size;\n \tenum object_type type;\n \tstruct unpack_entry_stack_ent small_delta_stack[UNPACK_ENTRY_STACK_PREALLOC];\n \tstruct unpack_entry_stack_ent *delta_stack = small_delta_stack;\n@@ -1943,8 +1947,10 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n \t\t\t      (uintmax_t)curpos, p->pack_name);\n \t\t\tdata = NULL;\n \t\t} else {\n+\t\t\tunsigned long sz;\n \t\t\tdata = patch_delta(base, base_size, delta_data,\n-\t\t\t\t\t   delta_size, &size);\n+\t\t\t\t\t   delta_size, &sz);\n+\t\t\tsize = sz;\n \n \t\t\t/*\n \t\t\t * We could not apply the delta; warn the user, but\n@@ -2929,7 +2935,7 @@ int packfile_read_object_stream(struct odb_read_stream **out,\n \tstruct odb_packed_read_stream *stream;\n \tstruct pack_window *window = NULL;\n \tenum object_type in_pack_type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \tin_pack_type = unpack_object_header(pack, &window, &offset, &size);\n \tunuse_pack(&window);\ndiff --git a/packfile.h b/packfile.h\nindex 9b647da7dd..49d6bdecf6 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -456,9 +456,9 @@ off_t find_pack_entry_one(const struct object_id *oid, struct packed_git *);\n \n int is_pack_valid(struct packed_git *);\n void *unpack_entry(struct repository *r, struct packed_git *, off_t, enum object_type *, unsigned long *);\n-unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n+unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, size_t *sizep);\n unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n-int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\n+int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, size_t *);\n off_t get_delta_base(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t *curpos, enum object_type type,\n \t\t     off_t delta_obj_offset);\n-- \ngitgitgadget\n\n"},{"id":"542439","messageId":"3274cba862ae42a6813710410274a692ec0f5d29.1777393580.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH 4/6] delta, packfile: use size_t for delta header sizes","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:18Z","receivedAt":"2026-04-28T16:26:29Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe delta header decoding functions return unsigned long, which\ntruncates on Windows for objects larger than 4GB. Introduce size_t\nvariants get_delta_hdr_size_sz() and get_size_from_delta_sz() that\npreserve the full 64-bit size, and use them in packed_object_info()\nwhere the size is needed for streaming decisions.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n delta.h    | 14 ++++++++++++--\n packfile.c | 33 ++++++++++++++++++++++++---------\n 2 files changed, 36 insertions(+), 11 deletions(-)\n\ndiff --git a/delta.h b/delta.h\nindex 8a56ec0799..fad68cfc45 100644\n--- a/delta.h\n+++ b/delta.h\n@@ -86,8 +86,11 @@ void *patch_delta(const void *src_buf, unsigned long src_size,\n  * This must be called twice on the delta data buffer, first to get the\n  * expected source buffer size, and again to get the target buffer size.\n  */\n-static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n-\t\t\t\t\t       const unsigned char *top)\n+/*\n+ * Size_t variant that doesn't truncate - use for >4GB objects on Windows.\n+ */\n+static inline size_t get_delta_hdr_size_sz(const unsigned char **datap,\n+\t\t\t\t\t   const unsigned char *top)\n {\n \tconst unsigned char *data = *datap;\n \tsize_t cmd, size = 0;\n@@ -98,6 +101,13 @@ static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n \t\ti += 7;\n \t} while (cmd & 0x80 && data < top);\n \t*datap = data;\n+\treturn size;\n+}\n+\n+static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n+\t\t\t\t\t       const unsigned char *top)\n+{\n+\tsize_t size = get_delta_hdr_size_sz(datap, top);\n \treturn cast_size_t_to_ulong(size);\n }\n \ndiff --git a/packfile.c b/packfile.c\nindex fdae91dd11..4208f53046 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1161,9 +1161,12 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \treturn used;\n }\n \n-unsigned long get_size_from_delta(struct packed_git *p,\n-\t\t\t\t  struct pack_window **w_curs,\n-\t\t\t\t  off_t curpos)\n+/*\n+ * Size_t variant for >4GB delta results on Windows.\n+ */\n+static size_t get_size_from_delta_sz(struct packed_git *p,\n+\t\t\t\t     struct pack_window **w_curs,\n+\t\t\t\t     off_t curpos)\n {\n \tconst unsigned char *data;\n \tunsigned char delta_head[20], *in;\n@@ -1210,10 +1213,18 @@ unsigned long get_size_from_delta(struct packed_git *p,\n \tdata = delta_head;\n \n \t/* ignore base size */\n-\tget_delta_hdr_size(&data, delta_head+sizeof(delta_head));\n+\tget_delta_hdr_size_sz(&data, delta_head+sizeof(delta_head));\n \n \t/* Read the result size */\n-\treturn get_delta_hdr_size(&data, delta_head+sizeof(delta_head));\n+\treturn get_delta_hdr_size_sz(&data, delta_head+sizeof(delta_head));\n+}\n+\n+unsigned long get_size_from_delta(struct packed_git *p,\n+\t\t\t\t  struct pack_window **w_curs,\n+\t\t\t\t  off_t curpos)\n+{\n+\tsize_t size = get_size_from_delta_sz(p, w_curs, curpos);\n+\treturn cast_size_t_to_ulong(size);\n }\n \n int unpack_object_header(struct packed_git *p,\n@@ -1618,14 +1629,18 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n \t\t\t\tret = -1;\n \t\t\t\tgoto out;\n \t\t\t}\n-\t\t\t*oi->sizep = get_size_from_delta(p, &w_curs, tmp_pos);\n-\t\t\tif (*oi->sizep == 0) {\n+\t\t\t/*\n+\t\t\t * Use size_t variant to avoid die() on >4GB deltas.\n+\t\t\t * oi->sizep is unsigned long, so truncation may occur,\n+\t\t\t * but streaming code uses its own size_t tracking.\n+\t\t\t */\n+\t\t\tsize = get_size_from_delta_sz(p, &w_curs, tmp_pos);\n+\t\t\tif (size == 0) {\n \t\t\t\tret = -1;\n \t\t\t\tgoto out;\n \t\t\t}\n-\t\t} else {\n-\t\t\t*oi->sizep = size;\n \t\t}\n+\t\t*oi->sizep = (unsigned long)size;\n \t}\n \n \tif (oi->disk_sizep || (oi->mtimep && p->is_cruft)) {\n-- \ngitgitgadget\n\n"},{"id":"542440","messageId":"afa74a3a2b9caf9989055a9311309f590729d6c1.1777393580.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH 5/6] test-tool: add a helper to synthesize large packfiles","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:19Z","receivedAt":"2026-04-28T16:26:31Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nTo test Git's behavior with very large pack files, we need a way to\ngenerate such files quickly.\n\nA naive approach using only readily-available Git commands would take\nover 10 hours for a 4GB pack file, which is prohibitive.\n\nSide-stepping Git's machinery and actual zlib compression by writing\nuncompressed content with the appropriate zlib header makes things\nmuch faster. The fastest method using this approach generates many\nsmall, unreachable blob objects and takes about 1.5 minutes for 4GB.\nHowever, this cannot be used because we need to test git clone, which\nrequires a reachable commit history.\n\nGenerating many reachable commits with small, uncompressed blobs takes\nabout 4 minutes for 4GB. But this approach 1) does not reproduce the\nissues we want to fix (which require individual objects larger than\n4GB) and 2) is comparatively slow because of the many SHA-1\ncalculations.\n\nThe approach taken here generates a single large blob (filled with NUL\nbytes), along with the trees and commits needed to make it reachable.\nThis takes about 2.5 minutes for 4.5GB, which is the fastest option\nthat produces a valid, clonable repository with an object large enough\nto trigger the bugs we want to test.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Makefile                   |   1 +\n compat/zlib-compat.h       |   2 +\n t/helper/meson.build       |   1 +\n t/helper/test-synthesize.c | 250 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c       |   1 +\n t/helper/test-tool.h       |   1 +\n 6 files changed, 256 insertions(+)\n create mode 100644 t/helper/test-synthesize.c\n\ndiff --git a/Makefile b/Makefile\nindex cedc234173..85405cb5b8 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -872,6 +872,7 @@ TEST_BUILTINS_OBJS += test-submodule-config.o\n TEST_BUILTINS_OBJS += test-submodule-nested-repo-config.o\n TEST_BUILTINS_OBJS += test-submodule.o\n TEST_BUILTINS_OBJS += test-subprocess.o\n+TEST_BUILTINS_OBJS += test-synthesize.o\n TEST_BUILTINS_OBJS += test-trace2.o\n TEST_BUILTINS_OBJS += test-truncate.o\n TEST_BUILTINS_OBJS += test-userdiff.o\ndiff --git a/compat/zlib-compat.h b/compat/zlib-compat.h\nindex ac08276622..5078c5ef6c 100644\n--- a/compat/zlib-compat.h\n+++ b/compat/zlib-compat.h\n@@ -7,6 +7,8 @@\n # define z_stream_s zng_stream_s\n # define gz_header_s zng_gz_header_s\n \n+# define adler32(adler, buf, len) zng_adler32(adler, buf, len)\n+\n # define crc32(crc, buf, len) zng_crc32(crc, buf, len)\n \n # define inflate(strm, bits) zng_inflate(strm, bits)\ndiff --git a/t/helper/meson.build b/t/helper/meson.build\nindex 675e64c010..3235f10ab8 100644\n--- a/t/helper/meson.build\n+++ b/t/helper/meson.build\n@@ -69,6 +69,7 @@ test_tool_sources = [\n   'test-submodule-nested-repo-config.c',\n   'test-submodule.c',\n   'test-subprocess.c',\n+  'test-synthesize.c',\n   'test-tool.c',\n   'test-trace2.c',\n   'test-truncate.c',\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nnew file mode 100644\nindex 0000000000..3ce7078078\n--- /dev/null\n+++ b/t/helper/test-synthesize.c\n@@ -0,0 +1,250 @@\n+#define USE_THE_REPOSITORY_VARIABLE\n+\n+#include \"test-tool.h\"\n+#include \"git-compat-util.h\"\n+#include \"git-zlib.h\"\n+#include \"hash.h\"\n+#include \"hex.h\"\n+#include \"object-file.h\"\n+#include \"object.h\"\n+#include \"pack.h\"\n+#include \"parse-options.h\"\n+#include \"parse.h\"\n+#include \"repository.h\"\n+#include \"setup.h\"\n+#include \"strbuf.h\"\n+#include \"write-or-die.h\"\n+\n+#define BLOCK_SIZE 0xffff\n+static const unsigned char zeros[BLOCK_SIZE];\n+\n+/*\n+ * Write data as an uncompressed zlib stream.\n+ * For data larger than 64KB, writes multiple uncompressed blocks.\n+ * If data is NULL, writes zeros.\n+ * Updates the pack checksum context.\n+ */\n+static void write_uncompressed_zlib(FILE *f, struct git_hash_ctx *pack_ctx,\n+\t\t\t\t    const void *data, size_t len,\n+\t\t\t\t    const struct git_hash_algo *algo)\n+{\n+\tunsigned char zlib_header[2] = { 0x78, 0x01 }; /* CMF, FLG */\n+\tunsigned char block_header[5];\n+\tconst unsigned char *p = data;\n+\tsize_t remaining = len;\n+\tuint32_t adler = 1L; /* adler32 initial value */\n+\tunsigned char adler_buf[4];\n+\n+\t/* Write zlib header */\n+\tfwrite_or_die(f, zlib_header, sizeof(zlib_header));\n+\talgo->update_fn(pack_ctx, zlib_header, 2);\n+\n+\t/* Write uncompressed blocks (max 64KB each) */\n+\tdo {\n+\t\tsize_t block_len = remaining > BLOCK_SIZE ? BLOCK_SIZE : remaining;\n+\t\tint is_final = (block_len == remaining);\n+\t\tconst unsigned char *block_data = data ? p : zeros;\n+\n+\t\tblock_header[0] = is_final ? 0x01 : 0x00;\n+\t\tblock_header[1] = block_len & 0xff;\n+\t\tblock_header[2] = (block_len >> 8) & 0xff;\n+\t\tblock_header[3] = block_header[1] ^ 0xff;\n+\t\tblock_header[4] = block_header[2] ^ 0xff;\n+\n+\t\tfwrite_or_die(f, block_header, sizeof(block_header));\n+\t\talgo->update_fn(pack_ctx, block_header, 5);\n+\n+\t\tif (block_len) {\n+\t\t\tfwrite_or_die(f, block_data, block_len);\n+\t\t\talgo->update_fn(pack_ctx, block_data, block_len);\n+\t\t\tadler = adler32(adler, block_data, block_len);\n+\t\t}\n+\n+\t\tif (data)\n+\t\t\tp += block_len;\n+\t\tremaining -= block_len;\n+\t} while (remaining > 0);\n+\n+\t/* Write adler32 checksum */\n+\tput_be32(adler_buf, adler);\n+\tfwrite_or_die(f, adler_buf, sizeof(adler_buf));\n+\talgo->update_fn(pack_ctx, adler_buf, 4);\n+}\n+\n+/*\n+ * Write an uncompressed object to the pack file.\n+ * If `data == NULL`, it is treated like a buffer to NUL bytes.\n+ * Updates the pack checksum context.\n+ */\n+static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n+\t\t\t      enum object_type type,\n+\t\t\t      const void *data, size_t len,\n+\t\t\t      struct object_id *oid,\n+\t\t\t      const struct git_hash_algo *algo)\n+{\n+\tunsigned char pack_header[MAX_PACK_OBJECT_HEADER];\n+\tchar object_header[32];\n+\tint pack_header_len, object_header_len;\n+\tstruct git_hash_ctx ctx;\n+\n+\t/* Write pack object header */\n+\tpack_header_len = encode_in_pack_object_header(pack_header,\n+\t\t\t\t\t\t       sizeof(pack_header),\n+\t\t\t\t\t\t       type, len);\n+\tfwrite_or_die(f, pack_header, pack_header_len);\n+\talgo->update_fn(pack_ctx, pack_header, pack_header_len);\n+\n+\t/* Write the data as uncompressed zlib */\n+\twrite_uncompressed_zlib(f, pack_ctx, data, len, algo);\n+\n+\talgo->init_fn(&ctx);\n+\tobject_header_len = format_object_header(object_header,\n+\t\t\t\t\t\t sizeof(object_header),\n+\t\t\t\t\t\t type, len);\n+\talgo->update_fn(&ctx, object_header, object_header_len);\n+\tif (data)\n+\t\talgo->update_fn(&ctx, data, len);\n+\telse {\n+\t\tfor (size_t i = len / BLOCK_SIZE; i; i--)\n+\t\t\talgo->update_fn(&ctx, zeros, BLOCK_SIZE);\n+\t\talgo->update_fn(&ctx, zeros, len % BLOCK_SIZE);\n+\t}\n+\talgo->final_oid_fn(oid, &ctx);\n+}\n+\n+/*\n+ * Generate a pack file with a single large (>4GB) reachable object.\n+ *\n+ * Creates:\n+ *   1. A large blob (all NUL bytes)\n+ *   2. A tree containing that blob as \"file\"\n+ *   3. A commit using that tree\n+ *   4. The empty tree\n+ *   5. A child commit using the empty tree\n+ *\n+ * This is useful for testing that Git can handle objects larger than 4GB.\n+ */\n+static int generate_pack_with_large_object(const char *path, size_t blob_size,\n+\t\t\t\t\t   const struct git_hash_algo *algo)\n+{\n+\tFILE *f = xfopen(path, \"wb\");\n+\tstruct git_hash_ctx pack_ctx;\n+\tunsigned char pack_hash[GIT_MAX_RAWSZ];\n+\tstruct object_id blob_oid, tree_oid, commit_oid, empty_tree_oid, final_commit_oid;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tconst uint32_t object_count = 5;\n+\tstruct pack_header pack_header = {\n+\t\t.hdr_signature = htonl(PACK_SIGNATURE),\n+\t\t.hdr_version = htonl(PACK_VERSION),\n+\t\t.hdr_entries = htonl(object_count),\n+\t};\n+\n+\talgo->init_fn(&pack_ctx);\n+\n+\t/* Write pack header */\n+\tfwrite_or_die(f, &pack_header, sizeof(pack_header));\n+\talgo->update_fn(&pack_ctx, &pack_header, sizeof(pack_header));\n+\n+\t/* 1. Write the large blob */\n+\twrite_pack_object(f, &pack_ctx, OBJ_BLOB, NULL, blob_size, &blob_oid, algo);\n+\n+\t/* 2. Write tree containing the blob as \"file\" */\n+\tstrbuf_addf(&buf, \"100644 file%c\", '\\0');\n+\tstrbuf_add(&buf, blob_oid.hash, algo->rawsz);\n+\twrite_pack_object(f, &pack_ctx, OBJ_TREE, buf.buf, buf.len, &tree_oid, algo);\n+\n+\t/* 3. Write commit using that tree */\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf,\n+\t\t    \"tree %s\\n\"\n+\t\t    \"author A U Thor <author@example.com> 1234567890 +0000\\n\"\n+\t\t    \"committer C O Mitter <committer@example.com> 1234567890 +0000\\n\"\n+\t\t    \"\\n\"\n+\t\t    \"Large blob commit\\n\",\n+\t\t    oid_to_hex(&tree_oid));\n+\twrite_pack_object(f, &pack_ctx, OBJ_COMMIT, buf.buf, buf.len, &commit_oid, algo);\n+\n+\t/* 4. Write the empty tree */\n+\twrite_pack_object(f, &pack_ctx, OBJ_TREE, \"\", 0, &empty_tree_oid, algo);\n+\n+\t/* 5. Write final commit using empty tree, with previous commit as parent */\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf,\n+\t\t    \"tree %s\\n\"\n+\t\t    \"parent %s\\n\"\n+\t\t    \"author A U Thor <author@example.com> 1234567890 +0000\\n\"\n+\t\t    \"committer C O Mitter <committer@example.com> 1234567890 +0000\\n\"\n+\t\t    \"\\n\"\n+\t\t    \"Empty tree commit\\n\",\n+\t\t    oid_to_hex(&empty_tree_oid),\n+\t\t    oid_to_hex(&commit_oid));\n+\twrite_pack_object(f, &pack_ctx, OBJ_COMMIT, buf.buf, buf.len, &final_commit_oid, algo);\n+\n+\t/* Write pack trailer (checksum) */\n+\talgo->final_fn(pack_hash, &pack_ctx);\n+\tfwrite_or_die(f, pack_hash, algo->rawsz);\n+\tif (fclose(f))\n+\t\tdie_errno(_(\"could not close '%s'\"), path);\n+\n+\tstrbuf_release(&buf);\n+\n+\t/* Print the final commit OID so caller can set up refs */\n+\tprintf(\"%s\\n\", oid_to_hex(&final_commit_oid));\n+\n+\treturn 0;\n+}\n+\n+static int cmd__synthesize__pack(int argc, const char **argv,\n+\t\t\t\t const char *prefix UNUSED,\n+\t\t\t\t struct repository *repo)\n+{\n+\tint non_git;\n+\tint reachable_large = 0;\n+\tconst struct git_hash_algo *algo;\n+\tsize_t blob_size;\n+\tuintmax_t blob_size_u;\n+\tconst char *path;\n+\tconst char * const usage[] = {\n+\t\t\"test-tool synthesize pack \"\n+\t\t\"--reachable-large <blob-size> <filename>\",\n+\t\tNULL\n+\t};\n+\tstruct option options[] = {\n+\t\tOPT_BOOL(0, \"reachable-large\", &reachable_large,\n+\t\t\t N_(\"write a pack with a single reachable large blob\")),\n+\t\tOPT_END()\n+\t};\n+\n+\tsetup_git_directory_gently(&non_git);\n+\trepo = the_repository;\n+\talgo = repo->hash_algo;\n+\n+\targc = parse_options(argc, argv, NULL, options, usage,\n+\t\t\t     PARSE_OPT_KEEP_ARGV0);\n+\tif (argc != 3 || !reachable_large)\n+\t\tusage_with_options(usage, options);\n+\n+\tif (!git_parse_unsigned(argv[1], &blob_size_u,\n+\t\t\t\tmaximum_unsigned_value_of_type(size_t)))\n+\t\tdie(_(\"'%s' is not a valid blob size\"), argv[1]);\n+\tblob_size = blob_size_u;\n+\tpath = argv[2];\n+\n+\treturn !!generate_pack_with_large_object(path, blob_size, algo);\n+}\n+\n+int cmd__synthesize(int argc, const char **argv)\n+{\n+\tconst char *prefix = NULL;\n+\tchar const * const synthesize_usage[] = {\n+\t\t\"test-tool synthesize pack <options>\",\n+\t\tNULL,\n+\t};\n+\tparse_opt_subcommand_fn *fn = NULL;\n+\tstruct option options[] = {\n+\t\tOPT_SUBCOMMAND(\"pack\", &fn, cmd__synthesize__pack),\n+\t\tOPT_END()\n+\t};\n+\targc = parse_options(argc, argv, prefix, options, synthesize_usage, 0);\n+\treturn !!fn(argc, argv, prefix, NULL);\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex a7abc618b3..b71a22b43b 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -82,6 +82,7 @@ static struct test_cmd cmds[] = {\n \t{ \"submodule-config\", cmd__submodule_config },\n \t{ \"submodule-nested-repo-config\", cmd__submodule_nested_repo_config },\n \t{ \"subprocess\", cmd__subprocess },\n+\t{ \"synthesize\", cmd__synthesize },\n \t{ \"trace2\", cmd__trace2 },\n \t{ \"truncate\", cmd__truncate },\n \t{ \"userdiff\", cmd__userdiff },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 7f150fa1eb..f2885b33d5 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -75,6 +75,7 @@ int cmd__submodule(int argc, const char **argv);\n int cmd__submodule_config(int argc, const char **argv);\n int cmd__submodule_nested_repo_config(int argc, const char **argv);\n int cmd__subprocess(int argc, const char **argv);\n+int cmd__synthesize(int argc, const char **argv);\n int cmd__trace2(int argc, const char **argv);\n int cmd__truncate(int argc, const char **argv);\n int cmd__userdiff(int argc, const char **argv);\n-- \ngitgitgadget\n\n"},{"id":"542441","messageId":"a3019888d8465e0f77926a91a20db170fef6989d.1777393580.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH 6/6] t5608: add regression test for >4GB object clone","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T16:26:20Z","receivedAt":"2026-04-28T16:26:32Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe shift overflow bug in index-pack and unpack-objects caused incorrect\nobject size calculation when the encoded size required more than 32 bits\nof shift. This would result in corrupted or failed unpacking of objects\nlarger than 4GB.\n\nAdd a test that creates a pack file containing a 4GB+ blob using the\nnew 'test-tool synthesize pack --reachable-large' command, then clones\nthe repository to verify the fix works correctly.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t5608-clone-2gb.sh | 37 +++++++++++++++++++++++++++++++++++++\n 1 file changed, 37 insertions(+)\n\ndiff --git a/t/t5608-clone-2gb.sh b/t/t5608-clone-2gb.sh\nindex 87a8cd9f98..af93302dde 100755\n--- a/t/t5608-clone-2gb.sh\n+++ b/t/t5608-clone-2gb.sh\n@@ -49,4 +49,41 @@ test_expect_success 'clone - with worktree, file:// protocol' '\n \n '\n \n+test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n+\tlarge_blob_size=$((4*1024*1024*1024+1)) &&\n+\tgit init --bare 4gb-repo &&\n+\thead_oid=$(test-tool synthesize pack \\\n+\t\t--reachable-large \"$large_blob_size\" \\\n+\t\t4gb-repo/objects/pack/test.pack) &&\n+\tgit -C 4gb-repo index-pack objects/pack/test.pack &&\n+\tgit -C 4gb-repo update-ref refs/heads/main $head_oid &&\n+\tgit -C 4gb-repo symbolic-ref HEAD refs/heads/main\n+'\n+\n+test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n+\t# The synthesized pack has five objects, so a large unpack limit keeps\n+\t# fetch-pack on the unpack-objects path.\n+\tgit -c fetch.unpackLimit=100 clone --bare \\\n+\t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-unpack &&\n+\n+\t# Verify the large blob survived the clone by comparing its OID\n+\t# between source and clone.  We cannot use \"cat-file -s\" because\n+\t# object_info.sizep is still unsigned long, which truncates >4GB\n+\t# sizes on Windows.  OID equality proves content integrity since\n+\t# the clone already verified checksums via index-pack/unpack-objects.\n+\tsource_blob=$(git -C 4gb-repo rev-parse main^:file) &&\n+\tclone_blob=$(git -C 4gb-clone-unpack rev-parse main^:file) &&\n+\ttest \"$source_blob\" = \"$clone_blob\"\n+'\n+\n+test_expect_success SIZE_T_IS_64BIT 'clone with >4GB object via index-pack' '\n+\t# Force fetch-pack to hand the pack to index-pack instead.\n+\tgit -c fetch.unpackLimit=1 clone --bare \\\n+\t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-index &&\n+\n+\tsource_blob=$(git -C 4gb-repo rev-parse main^:file) &&\n+\tclone_blob=$(git -C 4gb-clone-index rev-parse main^:file) &&\n+\ttest \"$source_blob\" = \"$clone_blob\"\n+'\n+\n test_done\n-- \ngitgitgadget\n"},{"id":"542474","messageId":"1b2ce8fe-7c99-4dde-ae07-1443d03cf523@gmail.com","threadId":"65563","inReplyTo":"3274cba862ae42a6813710410274a692ec0f5d29.1777393580.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 4/6] delta, packfile: use size_t for delta header sizes","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-04-29T13:28:45Z","receivedAt":"2026-04-29T13:28:47Z","isPatch":true,"body":"On 4/28/26 12:26 PM, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> The delta header decoding functions return unsigned long, which\n> truncates on Windows for objects larger than 4GB. Introduce size_t\n> variants get_delta_hdr_size_sz() and get_size_from_delta_sz() that\n> preserve the full 64-bit size, and use them in packed_object_info()\n> where the size is needed for streaming decisions.\n\n> + * Size_t variant that doesn't truncate - use for >4GB objects on Windows.\n> + */\n> +static inline size_t get_delta_hdr_size_sz(const unsigned char **datap,\n> +\t\t\t\t\t   const unsigned char *top)\n...\n> +static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n> +\t\t\t\t\t       const unsigned char *top)\n> +{\n> +\tsize_t size = get_delta_hdr_size_sz(datap, top);\n>   \treturn cast_size_t_to_ulong(size);\n>   }\n\nI like this trick to use the 64-bit implementation and only to\ndown-cast for API compatibility. This allows a more gradual\ntransition than if we replaced ulongs with size_ts everywhere\nat once.\n\nThanks,\n-Stolee\n\n"},{"id":"542475","messageId":"e1e8837f-7374-4079-ba87-ab95dd156e33@gmail.com","threadId":"65563","inReplyTo":"a3019888d8465e0f77926a91a20db170fef6989d.1777393580.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 6/6] t5608: add regression test for >4GB object clone","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-04-29T13:34:21Z","receivedAt":"2026-04-29T13:34:23Z","isPatch":true,"body":"On 4/28/26 12:26 PM, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> The shift overflow bug in index-pack and unpack-objects caused incorrect\n> object size calculation when the encoded size required more than 32 bits\n> of shift. This would result in corrupted or failed unpacking of objects\n> larger than 4GB.\n> \n> Add a test that creates a pack file containing a 4GB+ blob using the\n> new 'test-tool synthesize pack --reachable-large' command, then clones\n> the repository to verify the fix works correctly.\n\nAs mentioned in the previous patch, constructing this large packfile\ntakes ~4 minutes in CI pipelines. That's quite a lot to handle for\nevery CI run.\n\n> +test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n\nYour prereq here prevents it from running on 32-bit builds, which is\ngood. However, I wonder if it would be worth also specifying these\ntests as expensive. It's less likely that these layers will be touched\noften, so it should be enough to run these on major occasions, such as\ntesting a release candidate.\n\n(I also think it's appropriate to have these tests _not_ be marked\nexpensive in their original contribution to git-for-windows/git,\nbecause the pull request build should prove that the tests work. And\nmaybe git-for-windows/git should keep them for every PR since that's\nwhere the tests matter the most.)\n\nI suppose this also is a question for Junio and our process for\nvalidating releases. Do we have a certain cadence where we run the\nexpensive tests? What has been our threshold for hiding a test case\nbehind the expensive label?\n\nThanks,\n-Stolee\n\n"},{"id":"542476","messageId":"7bed56a9-520c-4d7b-a4d6-f03d89667993@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/6] Handle cloning of objects larger than 4GB on Windows","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-04-29T13:35:54Z","receivedAt":"2026-04-29T13:35:55Z","isPatch":true,"body":"On 4/28/26 12:26 PM, Johannes Schindelin via GitGitGadget wrote:\n> On Windows, unsigned long is 32-bit even on 64-bit systems. This causes\n> multiple problems when Git handles objects larger than 4GB. This patch\n> series is a very targeted fix for a very early part of the problem: it\n> addresses the most fundamental truncation points that prevent a >4GB object\n> from surviving a clone at all.\n> \n> Specifically, this fixes:\n> \n>   * zlib's uLong wrapping and triggering BUG() assertions in the git_zstream\n>     wrapper\n>   * Object sizes being truncated in pack streaming, delta headers, and\n>     index-pack/unpack-objects\n>   * pack-objects re-encoding reused pack entries with a truncated size,\n>     producing corrupt packs on the wire\n\nI'm glad to see this progress in this direction. It's a big step!\n\n> Many other code paths still use unsigned long for object sizes (e.g.,\n> cat-file -s, object_info.sizep, the delta machinery) and will need their own\n> conversions. This series does not attempt to fix those.\n\nI appreciate the mechanisms used to keep the scope of change minimal. This\ndiffers from the typical \"replace all 'unsigned long's with 'size_t'\"\nproposals in some clever ways.\n\n> Based on work by @LordKiRon in git-for-windows/git#6076.\n> \n> The last two commits add a test helper that synthesizes a pack with a >4GB\n> blob and regression tests that clone it via both the unpack-objects and\n> index-pack code paths using file:// transport.\n\nMy biggest concern here is about how expensive this test is. I wonder if\nwe should mark it as expensive for the core project and leave it enabled by\ndefault in git-for-windows/git.\n\nThanks,\n-Stolee\n\n"},{"id":"542533","messageId":"20260430141320.GA6659@tb-raspi4","threadId":"65563","inReplyTo":"dc660106ea8511e6adc44d2b70e9a4ae8b18090e.1777393580.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/6] index-pack, unpack-objects: use size_t for object size","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-04-30T14:13:20Z","receivedAt":"2026-04-30T14:13:28Z","isPatch":true,"body":"On Tue, Apr 28, 2026 at 04:26:15PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> When unpacking objects from a packfile, the object size is decoded\n> from a variable-length encoding. On platforms where unsigned long is\n> 32-bit (such as Windows, even in 64-bit builds), the shift operation\n> overflows when decoding sizes larger than 4GB. The result is a\n> truncated size value, causing the unpacked object to be corrupted or\n> rejected.\n> \n> Fix this by changing the size variable to size_t, which is 64-bit on\n> 64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\n> space.\n> \n> This was originally authored by LordKiRon <https://github.com/LordKiRon>,\n> who preferred not to reveal their real name and therefore agreed that I\n> take over authorship.\n\nGood to see things moving forward.\n\nSee even\nhttps://github.com/git-for-windows/git/pull/2179\nwhich is probably obsolete soon.\n"},{"id":"542540","messageId":"20260501063805.GA2038915@coredump.intra.peff.net","threadId":"65563","inReplyTo":"e1e8837f-7374-4079-ba87-ab95dd156e33@gmail.com","subject":"Re: [PATCH 6/6] t5608: add regression test for >4GB object clone","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-05-01T06:38:05Z","receivedAt":"2026-05-01T06:38:12Z","isPatch":true,"body":"On Wed, Apr 29, 2026 at 09:34:21AM -0400, Derrick Stolee wrote:\n\n> As mentioned in the previous patch, constructing this large packfile\n> takes ~4 minutes in CI pipelines. That's quite a lot to handle for\n> every CI run.\n\nAnd for local runs, too. ;) The test suite takes less than 90 seconds to\nrun on my laptop, but t5608 by itself 160 seconds. Even if running it in\nparallel didn't slow down the rest of the suite (which is not true,\nbecause it's obviously hogging a whole processor the whole time), that's\nstill almost doubling the run-time.\n\nI'd also worry about assuming that the trash directory can hold 4+GB\n(maybe 8GB+ since we clone it?), especially since many of us use ram\ndisks.\n\nThat said...\n\n> > +test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n> \n> Your prereq here prevents it from running on 32-bit builds, which is\n> good. However, I wonder if it would be worth also specifying these\n> tests as expensive. It's less likely that these layers will be touched\n> often, so it should be enough to run these on major occasions, such as\n> testing a release candidate.\n\nI think it is already skipped in most cases, because t5608 requires the\nGIT_TEST_CLONE_2GB environment variable be set. Arguably it should just\nbe using EXPENSIVE, too, as I do not think there is much value in having\nindividual flags for all of the expensive tests. I think that test just\npredates the modern prereq system entirely.\n\n> I suppose this also is a question for Junio and our process for\n> validating releases. Do we have a certain cadence where we run the\n> expensive tests? What has been our threshold for hiding a test case\n> behind the expensive label?\n\nAFAIK the labeling of expensive things is mostly ad-hoc, and nobody is\nsystematically running them. Likewise for the t/perf tests, which are\nsuper expensive but do (very occasionally) turn up interesting\nregressions.\n\n-Peff\n"},{"id":"542548","messageId":"ed226571-a095-456b-9d9e-bcc545d7ddfb@gmail.com","threadId":"65563","inReplyTo":"20260501063805.GA2038915@coredump.intra.peff.net","subject":"Re: [PATCH 6/6] t5608: add regression test for >4GB object clone","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-05-01T13:19:14Z","receivedAt":"2026-05-01T13:19:16Z","isPatch":true,"body":"On 5/1/2026 2:38 AM, Jeff King wrote:\n> On Wed, Apr 29, 2026 at 09:34:21AM -0400, Derrick Stolee wrote:\n\n>>> +test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n>>\n>> Your prereq here prevents it from running on 32-bit builds, which is\n>> good. However, I wonder if it would be worth also specifying these\n>> tests as expensive. It's less likely that these layers will be touched\n>> often, so it should be enough to run these on major occasions, such as\n>> testing a release candidate.\n> \n> I think it is already skipped in most cases, because t5608 requires the\n> GIT_TEST_CLONE_2GB environment variable be set. Arguably it should just\n> be using EXPENSIVE, too, as I do not think there is much value in having\n> individual flags for all of the expensive tests. I think that test just\n> predates the modern prereq system entirely.\n\nThanks for the extra details here! That helps avoid the issues that I\nwas thinking about, but maybe doubling-down and adding EXPENSIVE is\nstill worth it. \n>> I suppose this also is a question for Junio and our process for\n>> validating releases. Do we have a certain cadence where we run the\n>> expensive tests? What has been our threshold for hiding a test case\n>> behind the expensive label?\n> \n> AFAIK the labeling of expensive things is mostly ad-hoc, and nobody is\n> systematically running them. Likewise for the t/perf tests, which are\n> super expensive but do (very occasionally) turn up interesting\n> regressions.\nI used to be more diligent about running the performance tests myself\naround release windows. The EXPENSIVE tests would also be good to do\non rc0. I will contemplate how to put this into my routine.\n\nThanks,\n-Stolee\n\n"},{"id":"542609","messageId":"1d2034e3-5a1c-81e7-88a6-65558da2360a@gmx.de","threadId":"65563","inReplyTo":"20260430141320.GA6659@tb-raspi4","subject":"Re: [PATCH 1/6] index-pack, unpack-objects: use size_t for object size","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-03T14:46:45Z","receivedAt":"2026-05-03T14:46:48Z","isPatch":true,"body":"Hi Torsten,\n\nOn Sun, 3 May 2026, Torsten Bögershausen wrote:\n\n> On Tue, Apr 28, 2026 at 04:26:15PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > \n> > When unpacking objects from a packfile, the object size is decoded\n> > from a variable-length encoding. On platforms where unsigned long is\n> > 32-bit (such as Windows, even in 64-bit builds), the shift operation\n> > overflows when decoding sizes larger than 4GB. The result is a\n> > truncated size value, causing the unpacked object to be corrupted or\n> > rejected.\n> > \n> > Fix this by changing the size variable to size_t, which is 64-bit on\n> > 64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\n> > space.\n> > \n> > This was originally authored by LordKiRon\n> > <https://github.com/LordKiRon>, who preferred not to reveal their real\n> > name and therefore agreed that I take over authorship.\n> \n> Good to see things moving forward.\n> \n> See even\n> https://github.com/git-for-windows/git/pull/2179\n> which is probably obsolete soon.\n\nThe word \"probably\" is maybe a bit overwhelmed in this sentence by the\nsheer extent of what still needs to be done.\n\nThere have been multiple contributors who despaired over the task [*1*] to\nsplit this PR apart in ways that would stand a chance to be accepted (or\nfor that matter: reviewed) on the Git mailing list...\n\nBut yes, it is my hope that I can wittle down that PR into more\ncontributions along the lines of f9ba6acaa934 (Merge branch\n'mc/clean-smudge-with-llp64', 2021-11-29). I still have to upstream\nhttps://github.com/git-for-windows/git/pull/3533, which is another teeny\ntiny step toward completing what the PR you mentioned set out to\naccomplish.\n\nAnd yes, the pattern that Stolee already recognized, to duplicate function\nsignatures into the ones that the 20th century kindly asked to be returned\nand the ones using `size_t` instead, this is the pattern I specifically\nwanted to use so that incremental patch series have a chance of getting\nreviews and to trickle into git/git.\n\nCiao,\nJohannes\n\nFootnote *1*: I specifically asked for this kind of splitting-out:\nhttps://github.com/git-for-windows/git/pull/2179#issuecomment-525926366\n"},{"id":"542610","messageId":"d40c5d8a-9f03-43fd-f6bf-0b7516d21804@gmx.de","threadId":"65563","inReplyTo":"1b2ce8fe-7c99-4dde-ae07-1443d03cf523@gmail.com","subject":"Re: [PATCH 4/6] delta, packfile: use size_t for delta header sizes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-03T14:49:19Z","receivedAt":"2026-05-03T14:49:22Z","isPatch":true,"body":"Hi Stolee,\n\nOn Sun, 3 May 2026, Derrick Stolee wrote:\n\n> On 4/28/26 12:26 PM, Johannes Schindelin via GitGitGadget wrote:\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > \n> > The delta header decoding functions return unsigned long, which\n> > truncates on Windows for objects larger than 4GB. Introduce size_t\n> > variants get_delta_hdr_size_sz() and get_size_from_delta_sz() that\n> > preserve the full 64-bit size, and use them in packed_object_info()\n> > where the size is needed for streaming decisions.\n> \n> > + * Size_t variant that doesn't truncate - use for >4GB objects on Windows.\n> > + */\n> > +static inline size_t get_delta_hdr_size_sz(const unsigned char **datap,\n> > +\t\t\t\t\t   const unsigned char *top)\n> ...\n> > +static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n> > +\t\t\t\t\t       const unsigned char *top)\n> > +{\n> > +\tsize_t size = get_delta_hdr_size_sz(datap, top);\n> >   \treturn cast_size_t_to_ulong(size);\n> >   }\n> \n> I like this trick to use the 64-bit implementation and only to\n> down-cast for API compatibility. This allows a more gradual\n> transition than if we replaced ulongs with size_ts everywhere\n> at once.\n\nThank you! That was indeed the exact thing I wanted to achieve: To allow\nfor incremental, easy-to-review patch series.\n\nCiao,\nJohannes\n"},{"id":"542683","messageId":"7ab69861-1508-106e-5d55-5ce5508653f8@gmx.de","threadId":"65563","inReplyTo":"ed226571-a095-456b-9d9e-bcc545d7ddfb@gmail.com","subject":"Re: [PATCH 6/6] t5608: add regression test for >4GB object clone","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-04T17:07:10Z","receivedAt":"2026-05-04T17:07:15Z","isPatch":true,"body":"Hi Stolee & Jeff,\n\nOn Sun, 3 May 2026, Derrick Stolee wrote:\n\n> On 5/1/2026 2:38 AM, Jeff King wrote:\n> > On Wed, Apr 29, 2026 at 09:34:21AM -0400, Derrick Stolee wrote:\n> \n> >>> +test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n> >>\n> >> Your prereq here prevents it from running on 32-bit builds, which is\n> >> good. However, I wonder if it would be worth also specifying these\n> >> tests as expensive. It's less likely that these layers will be touched\n> >> often, so it should be enough to run these on major occasions, such as\n> >> testing a release candidate.\n> > \n> > I think it is already skipped in most cases, because t5608 requires the\n> > GIT_TEST_CLONE_2GB environment variable be set. Arguably it should just\n> > be using EXPENSIVE, too, as I do not think there is much value in having\n> > individual flags for all of the expensive tests. I think that test just\n> > predates the modern prereq system entirely.\n> \n> Thanks for the extra details here! That helps avoid the issues that I\n> was thinking about, but maybe doubling-down and adding EXPENSIVE is\n> still worth it. \n\nIndeed. The `GIT_TEST_CLONE_2GB` flag is set for all CI jobs:\nhttps://gitlab.com/git-scm/git/-/blob/v2.54.0/ci/lib.sh#L314\n\nSo I've tried to accelerate it.\n\nFirst, by using the \"unsafe\" SHA-1 (because we're in no danger here, we\ngenerate the data ourselves). That helped some, but unfortunately the most\nefficient implementation (OpenSSL) is out of bounds for us because we\ncannot enable its use in the default configuration for licensing reasons:\nOpenSSL's license and Git's GPLv2 are fundamentally incompatible with each\nother, and even though Linus' intention with\nhttps://gitlab.com/git-scm/git/-/blob/e83c5163316f89bfbde7d9ab23ca2e25604af290/Makefile#L11\nwas clearly to allow linking to OpenSSL (actually, requiring it), there is\none Git contributor I spoke to who flatly stated that they'd block every\nattempt to add an exception to Git's license retroactively to allow\nlinking to OpenSSL (at least for distribution, which is when the GPLv2\nwould kick in). So that's that.\n\nSecond, I added shortcuts for the 4GiB+1 size exercised in the added test\ncases. That did help! But only the generation time of those packfiles is\nhelpd by that. The `git clone` that is tested is still awfully slow, and\nthat _cannot_ be worked around in the same ways.\n\nTherefore, I ended up marking the test cases as `EXPENSIVE`.\n\nTo allow them to be run regularly anyway, I added a final patch to the\nseries that lets the CI runs on the integration branches (other than\n`seen`) run all `EXPENSIVE` test cases. One could argue that this patch\nshould be split out, and I'm open to it, but in the interest of keeping\nthe time I am working on this patch series _somewhat_ closer to what could\nbe called reasonable, I'll just keep it in the patch series for now,\nhedging for the possibility that maybe nobody objects?\n\nCiao,\nJohannes\n\n> >> I suppose this also is a question for Junio and our process for\n> >> validating releases. Do we have a certain cadence where we run the\n> >> expensive tests? What has been our threshold for hiding a test case\n> >> behind the expensive label?\n> > \n> > AFAIK the labeling of expensive things is mostly ad-hoc, and nobody is\n> > systematically running them. Likewise for the t/perf tests, which are\n> > super expensive but do (very occasionally) turn up interesting\n> > regressions.\n> I used to be more diligent about running the performance tests myself\n> around release windows. The EXPENSIVE tests would also be good to do\n> on rc0. I will contemplate how to put this into my routine.\n> \n> Thanks,\n> -Stolee\n> \n> \n> \n"},{"id":"542684","messageId":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.git.1777393580.gitgitgadget@gmail.com","subject":"[PATCH v2 00/11] Handle cloning of objects larger than 4GB on Windows","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:17Z","receivedAt":"2026-05-04T17:08:31Z","isPatch":true,"body":"On Windows, unsigned long is 32-bit even on 64-bit systems. This causes\nmultiple problems when Git handles objects larger than 4GB. This patch\nseries is a very targeted fix for a very early part of the problem: it\naddresses the most fundamental truncation points that prevent a >4GB object\nfrom surviving a clone at all.\n\nSpecifically, this fixes:\n\n * zlib's uLong wrapping and triggering BUG() assertions in the git_zstream\n   wrapper\n * Object sizes being truncated in pack streaming, delta headers, and\n   index-pack/unpack-objects\n * pack-objects re-encoding reused pack entries with a truncated size,\n   producing corrupt packs on the wire\n\nMany other code paths still use unsigned long for object sizes (e.g.,\ncat-file -s, object_info.sizep, the delta machinery) and will need their own\nconversions. This series does not attempt to fix those.\n\nBased on work by @LordKiRon in git-for-windows/git#6076.\n\nThe last two commits add a test helper that synthesizes a pack with a >4GB\nblob and regression tests that clone it via both the unpack-objects and\nindex-pack code paths using file:// transport.\n\nChanges since v1:\n\n * dramatically accelerated the test helper that generates 4GB pack files,\n   via two separate strategies:\n   1. using the \"unsafe\" SHA-1 for the blob OID computation.\n   2. using pre-computed \"Lego blocks\" to construct the 4GB packs needed in\n      the test cases, where the size (and therefore the involved OIDs) are\n      well-known in advance.\n * even with these improvements, the actual git clone is still slow (of\n   course, because it cannot use any of those shortcuts), therefore the\n   tests are marked as EXPENSIVE.\n * to exercise those tests nevertheless, the last patch lets all EXPENSIVE\n   test cases be run for the integration branches other than seen.\n\nJohannes Schindelin (11):\n  index-pack, unpack-objects: use size_t for object size\n  git-zlib: handle data streams larger than 4GB\n  odb, packfile: use size_t for streaming object sizes\n  delta, packfile: use size_t for delta header sizes\n  test-tool: add a helper to synthesize large packfiles\n  t5608: add regression test for >4GB object clone\n  test-tool synthesize: use the unsafe hash for speed\n  test-tool synthesize: precompute pack for 4 GiB + 1\n  test-tool synthesize: add precomputed SHA-256 pack for 4 GiB + 1\n  t5608: mark >4GB tests as EXPENSIVE\n  ci: run expensive tests on push builds to integration branches\n\n Makefile                     |   1 +\n builtin/index-pack.c         |   9 +-\n builtin/pack-objects.c       |  23 +-\n builtin/unpack-objects.c     |   5 +-\n ci/lib.sh                    |   9 +\n compat/zlib-compat.h         |   2 +\n delta.h                      |  14 +-\n git-zlib.c                   |  25 +-\n git-zlib.h                   |   4 +-\n object-file.c                |  12 +-\n odb/streaming.c              |  13 +-\n odb/streaming.h              |   2 +-\n oss-fuzz/fuzz-pack-headers.c |   2 +-\n pack-bitmap.c                |   2 +-\n pack-check.c                 |   6 +-\n packfile.c                   |  57 ++--\n packfile.h                   |   4 +-\n t/helper/meson.build         |   1 +\n t/helper/test-synthesize.c   | 541 +++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c         |   1 +\n t/helper/test-tool.h         |   1 +\n t/t5608-clone-2gb.sh         |  37 +++\n 22 files changed, 718 insertions(+), 53 deletions(-)\n create mode 100644 t/helper/test-synthesize.c\n\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2102%2Fdscho%2Ffix-large-clones-on-windows-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2102/dscho/fix-large-clones-on-windows-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/2102\n\nRange-diff vs v1:\n\n  1:  dc660106ea =  1:  dc660106ea index-pack, unpack-objects: use size_t for object size\n  2:  92f4327b1f =  2:  92f4327b1f git-zlib: handle data streams larger than 4GB\n  3:  3a539061c5 =  3:  3a539061c5 odb, packfile: use size_t for streaming object sizes\n  4:  3274cba862 =  4:  3274cba862 delta, packfile: use size_t for delta header sizes\n  5:  afa74a3a2b =  5:  afa74a3a2b test-tool: add a helper to synthesize large packfiles\n  6:  a3019888d8 =  6:  a3019888d8 t5608: add regression test for >4GB object clone\n  -:  ---------- >  7:  859e93e7a9 test-tool synthesize: use the unsafe hash for speed\n  -:  ---------- >  8:  29b9a74e91 test-tool synthesize: precompute pack for 4 GiB + 1\n  -:  ---------- >  9:  8e6e720804 test-tool synthesize: add precomputed SHA-256 pack for 4 GiB + 1\n  -:  ---------- > 10:  5b44410b2f t5608: mark >4GB tests as EXPENSIVE\n  -:  ---------- > 11:  1eaaa7fad7 ci: run expensive tests on push builds to integration branches\n\n-- \ngitgitgadget\n"},{"id":"542685","messageId":"dc660106ea8511e6adc44d2b70e9a4ae8b18090e.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:18Z","receivedAt":"2026-05-04T17:08:32Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen unpacking objects from a packfile, the object size is decoded\nfrom a variable-length encoding. On platforms where unsigned long is\n32-bit (such as Windows, even in 64-bit builds), the shift operation\noverflows when decoding sizes larger than 4GB. The result is a\ntruncated size value, causing the unpacked object to be corrupted or\nrejected.\n\nFix this by changing the size variable to size_t, which is 64-bit on\n64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\nspace.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n builtin/index-pack.c     | 9 +++++----\n builtin/unpack-objects.c | 5 +++--\n 2 files changed, 8 insertions(+), 6 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..cc660582e9 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -37,7 +37,7 @@ static const char index_pack_usage[] =\n \n struct object_entry {\n \tstruct pack_idx_entry idx;\n-\tunsigned long size;\n+\tsize_t size;\n \tunsigned char hdr_size;\n \tsigned char type;\n \tsigned char real_type;\n@@ -469,7 +469,7 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n-static void *unpack_entry_data(off_t offset, unsigned long size,\n+static void *unpack_entry_data(off_t offset, size_t size,\n \t\t\t       enum object_type type, struct object_id *oid)\n {\n \tstatic char fixed_buf[8192];\n@@ -524,7 +524,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \t\t\t      struct object_id *oid)\n {\n \tunsigned char *p;\n-\tunsigned long size, c;\n+\tsize_t size;\n+\tunsigned long c;\n \toff_t base_offset;\n \tunsigned shift;\n \tvoid *data;\n@@ -542,7 +543,7 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \t\tp = fill(1);\n \t\tc = *p;\n \t\tuse(1);\n-\t\tsize += (c & 0x7f) << shift;\n+\t\tsize += ((size_t)c & 0x7f) << shift;\n \t\tshift += 7;\n \t}\n \tobj->size = size;\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..59a36c2481 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -533,7 +533,8 @@ static void unpack_one(unsigned nr)\n {\n \tunsigned shift;\n \tunsigned char *pack;\n-\tunsigned long size, c;\n+\tsize_t size;\n+\tunsigned long c;\n \tenum object_type type;\n \n \tobj_list[nr].offset = consumed_bytes;\n@@ -548,7 +549,7 @@ static void unpack_one(unsigned nr)\n \t\tpack = fill(1);\n \t\tc = *pack;\n \t\tuse(1);\n-\t\tsize += (c & 0x7f) << shift;\n+\t\tsize += ((size_t)c & 0x7f) << shift;\n \t\tshift += 7;\n \t}\n \n-- \ngitgitgadget\n\n"},{"id":"542686","messageId":"92f4327b1fe09126dd6421b071a9071ce5530371.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 02/11] git-zlib: handle data streams larger than 4GB","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:19Z","receivedAt":"2026-05-04T17:08:33Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOn Windows, zlib's `uLong` type is 32-bit even on 64-bit systems. When\nprocessing data streams larger than 4GB, the `total_in` and `total_out`\nfields in zlib's `z_stream` structure wrap around, which caused the\nsanity checks in `zlib_post_call()` to trigger `BUG()` assertions.\n\nThe git_zstream wrapper now tracks its own 64-bit totals rather than\ncopying them from zlib. The sanity checks compare only the low bits,\nusing `maximum_unsigned_value_of_type(uLong)` to mask appropriately for\nthe platform's `uLong` size.\n\nThis is based on work by LordKiRon in git-for-windows#6076.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n git-zlib.c    | 25 +++++++++++++++++--------\n git-zlib.h    |  4 ++--\n object-file.c |  2 +-\n 3 files changed, 20 insertions(+), 11 deletions(-)\n\ndiff --git a/git-zlib.c b/git-zlib.c\nindex df9604910e..b91cb323ae 100644\n--- a/git-zlib.c\n+++ b/git-zlib.c\n@@ -30,6 +30,9 @@ static const char *zerr_to_string(int status)\n  */\n /* #define ZLIB_BUF_MAX ((uInt)-1) */\n #define ZLIB_BUF_MAX ((uInt) 1024 * 1024 * 1024) /* 1GB */\n+\n+/* uLong is 32-bit on Windows, even on 64-bit systems */\n+#define ULONG_MAX_VALUE maximum_unsigned_value_of_type(uLong)\n static inline uInt zlib_buf_cap(unsigned long len)\n {\n \treturn (ZLIB_BUF_MAX < len) ? ZLIB_BUF_MAX : len;\n@@ -39,31 +42,37 @@ static void zlib_pre_call(git_zstream *s)\n {\n \ts->z.next_in = s->next_in;\n \ts->z.next_out = s->next_out;\n-\ts->z.total_in = s->total_in;\n-\ts->z.total_out = s->total_out;\n+\ts->z.total_in = (uLong)(s->total_in & ULONG_MAX_VALUE);\n+\ts->z.total_out = (uLong)(s->total_out & ULONG_MAX_VALUE);\n \ts->z.avail_in = zlib_buf_cap(s->avail_in);\n \ts->z.avail_out = zlib_buf_cap(s->avail_out);\n }\n \n static void zlib_post_call(git_zstream *s, int status)\n {\n-\tunsigned long bytes_consumed;\n-\tunsigned long bytes_produced;\n+\tsize_t bytes_consumed;\n+\tsize_t bytes_produced;\n \n \tbytes_consumed = s->z.next_in - s->next_in;\n \tbytes_produced = s->z.next_out - s->next_out;\n-\tif (s->z.total_out != s->total_out + bytes_produced)\n+\t/*\n+\t * zlib's total_out/total_in are uLong which may wrap for >4GB.\n+\t * We track our own totals and verify only the low bits match.\n+\t */\n+\tif ((s->z.total_out & ULONG_MAX_VALUE) !=\n+\t    ((s->total_out + bytes_produced) & ULONG_MAX_VALUE))\n \t\tBUG(\"total_out mismatch\");\n \t/*\n \t * zlib does not update total_in when it returns Z_NEED_DICT,\n \t * causing a mismatch here. Skip the sanity check in that case.\n \t */\n \tif (status != Z_NEED_DICT &&\n-\t    s->z.total_in != s->total_in + bytes_consumed)\n+\t    (s->z.total_in & ULONG_MAX_VALUE) !=\n+\t    ((s->total_in + bytes_consumed) & ULONG_MAX_VALUE))\n \t\tBUG(\"total_in mismatch\");\n \n-\ts->total_out = s->z.total_out;\n-\ts->total_in = s->z.total_in;\n+\ts->total_out += bytes_produced;\n+\ts->total_in += bytes_consumed;\n \t/* zlib-ng marks `next_in` as `const`, so we have to cast it away. */\n \ts->next_in = (unsigned char *) s->z.next_in;\n \ts->next_out = s->z.next_out;\ndiff --git a/git-zlib.h b/git-zlib.h\nindex 0e66fefa8c..44380e8ad3 100644\n--- a/git-zlib.h\n+++ b/git-zlib.h\n@@ -7,8 +7,8 @@ typedef struct git_zstream {\n \tstruct z_stream_s z;\n \tunsigned long avail_in;\n \tunsigned long avail_out;\n-\tunsigned long total_in;\n-\tunsigned long total_out;\n+\tsize_t total_in;\n+\tsize_t total_out;\n \tunsigned char *next_in;\n \tunsigned char *next_out;\n } git_zstream;\ndiff --git a/object-file.c b/object-file.c\nindex 2acc9522df..086b2b65ff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1118,7 +1118,7 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n \n \tif (stream.total_in != len + hdrlen)\n-\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\tdie(_(\"write stream object %\"PRIuMAX\" != %\"PRIuMAX), (uintmax_t)stream.total_in,\n \t\t    (uintmax_t)len + hdrlen);\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"542687","messageId":"3a539061c5f62c65d46bd0eb774bb1b1239463ff.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 03/11] odb, packfile: use size_t for streaming object sizes","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:20Z","receivedAt":"2026-05-04T17:08:34Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe odb_read_stream structure uses unsigned long for the size field,\nwhich is 32-bit on Windows even in 64-bit builds. When streaming\nobjects larger than 4GB, the size would be truncated to zero or an\nincorrect value, resulting in empty files being written to disk.\n\nChange the size field in odb_read_stream to size_t and introduce\nunpack_object_header_sz() to return sizes via size_t pointer. Since\nobject_info.sizep remains unsigned long for API compatibility, use\ntemporary variables where the types differ, with comments noting the\ntruncation limitation for code paths that still use unsigned long.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n builtin/pack-objects.c       | 23 ++++++++++++++++-------\n object-file.c                | 10 +++++++++-\n odb/streaming.c              | 13 ++++++++++++-\n odb/streaming.h              |  2 +-\n oss-fuzz/fuzz-pack-headers.c |  2 +-\n pack-bitmap.c                |  2 +-\n pack-check.c                 |  6 ++++--\n packfile.c                   | 24 +++++++++++++++---------\n packfile.h                   |  4 ++--\n 9 files changed, 61 insertions(+), 25 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex dd2480a73d..aa4b1cb9b8 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -629,14 +629,21 @@ static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n \tstruct packed_git *p = IN_PACK(entry);\n \tstruct pack_window *w_curs = NULL;\n \tuint32_t pos;\n-\toff_t offset;\n+\toff_t offset, cur;\n \tenum object_type type = oe_type(entry);\n+\tenum object_type in_pack_type;\n \toff_t datalen;\n \tunsigned char header[MAX_PACK_OBJECT_HEADER],\n \t\t      dheader[MAX_PACK_OBJECT_HEADER];\n \tunsigned hdrlen;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n-\tunsigned long entry_size = SIZE(entry);\n+\tsize_t entry_size;\n+\n+\tcur = entry->in_pack_offset;\n+\tin_pack_type = unpack_object_header(p, &w_curs, &cur, &entry_size);\n+\tif (in_pack_type < 0)\n+\t\tdie(_(\"write_reuse_object: unable to parse object header of %s\"),\n+\t\t    oid_to_hex(&entry->idx.oid));\n \n \tif (DELTA(entry))\n \t\ttype = (allow_ofs_delta && DELTA(entry)->idx.offset) ?\n@@ -1087,7 +1094,7 @@ static void write_reused_pack_one(struct packed_git *reuse_packfile,\n {\n \toff_t offset, next, cur;\n \tenum object_type type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \toffset = pack_pos_to_offset(reuse_packfile, pos);\n \tnext = pack_pos_to_offset(reuse_packfile, pos + 1);\n@@ -2243,7 +2250,7 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n \t\toff_t ofs;\n \t\tunsigned char *buf, c;\n \t\tenum object_type type;\n-\t\tunsigned long in_pack_size;\n+\t\tsize_t in_pack_size;\n \n \t\tbuf = use_pack(p, &w_curs, entry->in_pack_offset, &avail);\n \n@@ -2734,16 +2741,18 @@ unsigned long oe_get_size_slow(struct packing_data *pack,\n \tstruct pack_window *w_curs;\n \tunsigned char *buf;\n \tenum object_type type;\n-\tunsigned long used, avail, size;\n+\tunsigned long used, avail;\n+\tsize_t size;\n \n \tif (e->type_ != OBJ_OFS_DELTA && e->type_ != OBJ_REF_DELTA) {\n+\t\tunsigned long sz;\n \t\tpacking_data_lock(&to_pack);\n \t\tif (odb_read_object_info(the_repository->objects,\n-\t\t\t\t\t &e->idx.oid, &size) < 0)\n+\t\t\t\t\t &e->idx.oid, &sz) < 0)\n \t\t\tdie(_(\"unable to get size of %s\"),\n \t\t\t    oid_to_hex(&e->idx.oid));\n \t\tpacking_data_unlock(&to_pack);\n-\t\treturn size;\n+\t\treturn sz;\n \t}\n \n \tp = oe_in_pack(pack, e);\ndiff --git a/object-file.c b/object-file.c\nindex 086b2b65ff..0be2981c7a 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2326,6 +2326,7 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_loose_read_stream *st;\n \tunsigned long mapsize;\n+\tunsigned long size_ul;\n \tvoid *mapped;\n \n \tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n@@ -2349,11 +2350,18 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n \t\tgoto error;\n \t}\n \n-\toi.sizep = &st->base.size;\n+\t/*\n+\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n+\t * st->base.size is size_t (64-bit). Use temporary variable.\n+\t * Note: loose objects >4GB would still truncate here, but such\n+\t * large loose objects are uncommon (they'd normally be packed).\n+\t */\n+\toi.sizep = &size_ul;\n \toi.typep = &st->base.type;\n \n \tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n \t\tgoto error;\n+\tst->base.size = size_ul;\n \n \tst->mapped = mapped;\n \tst->mapsize = mapsize;\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 5927a12954..af2adf5ce7 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -157,15 +157,26 @@ static int open_istream_incore(struct odb_read_stream **out,\n \t\t.base.read = read_istream_incore,\n \t};\n \tstruct odb_incore_read_stream *st;\n+\tunsigned long size_ul;\n \tint ret;\n \n \toi.typep = &stream.base.type;\n-\toi.sizep = &stream.base.size;\n+\t/*\n+\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n+\t * stream.base.size is size_t (64-bit). We use a temporary variable\n+\t * because the types are incompatible. Note: this path still truncates\n+\t * for >4GB objects, but large objects should use pack streaming\n+\t * (packfile_store_read_object_stream) which handles size_t properly.\n+\t * This incore fallback is only used for small objects or when pack\n+\t * streaming is unavailable.\n+\t */\n+\toi.sizep = &size_ul;\n \toi.contentp = (void **)&stream.buf;\n \tret = odb_read_object_info_extended(odb, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n+\tstream.base.size = size_ul;\n \n \tCALLOC_ARRAY(st, 1);\n \t*st = stream;\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex c7861f7e13..517e2ea2d3 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -21,7 +21,7 @@ struct odb_read_stream {\n \todb_read_stream_close_fn close;\n \todb_read_stream_read_fn read;\n \tenum object_type type;\n-\tunsigned long size; /* inflated size of full object */\n+\tsize_t size; /* inflated size of full object */\n };\n \n /*\ndiff --git a/oss-fuzz/fuzz-pack-headers.c b/oss-fuzz/fuzz-pack-headers.c\nindex 150c0f5fa2..ef61ab577c 100644\n--- a/oss-fuzz/fuzz-pack-headers.c\n+++ b/oss-fuzz/fuzz-pack-headers.c\n@@ -6,7 +6,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size);\n int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n {\n \tenum object_type type;\n-\tunsigned long len;\n+\tsize_t len;\n \n \tunpack_object_header_buffer((const unsigned char *)data,\n \t\t\t\t    (unsigned long)size, &type, &len);\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex f6ec18d83a..f9af8a96bd 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -2270,7 +2270,7 @@ static int try_partial_reuse(struct bitmap_index *bitmap_git,\n {\n \toff_t delta_obj_offset;\n \tenum object_type type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \tif (pack_pos >= pack->p->num_objects)\n \t\treturn -1; /* not actually in the pack */\ndiff --git a/pack-check.c b/pack-check.c\nindex 79992bb509..2792f34d25 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -110,7 +110,7 @@ static int verify_packfile(struct repository *r,\n \t\tvoid *data;\n \t\tstruct object_id oid;\n \t\tenum object_type type;\n-\t\tunsigned long size;\n+\t\tsize_t size;\n \t\toff_t curpos;\n \t\tint data_valid;\n \n@@ -143,7 +143,9 @@ static int verify_packfile(struct repository *r,\n \t\t\tdata = NULL;\n \t\t\tdata_valid = 0;\n \t\t} else {\n-\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &size);\n+\t\t\tunsigned long sz;\n+\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &sz);\n+\t\t\tsize = sz;\n \t\t\tdata_valid = 1;\n \t\t}\n \ndiff --git a/packfile.c b/packfile.c\nindex b012d648ad..fdae91dd11 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1133,7 +1133,7 @@ out:\n }\n \n unsigned long unpack_object_header_buffer(const unsigned char *buf,\n-\t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n+\t\tunsigned long len, enum object_type *type, size_t *sizep)\n {\n \tunsigned shift;\n \tsize_t size, c;\n@@ -1144,7 +1144,11 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \tsize = c & 15;\n \tshift = 4;\n \twhile (c & 0x80) {\n-\t\tif (len <= used || (bitsizeof(long) - 7) < shift) {\n+\t\t/*\n+\t\t * Each continuation byte adds 7 bits. Ensure shift won't\n+\t\t * overflow size_t (use size_t not long for 64-bit on Windows).\n+\t\t */\n+\t\tif (len <= used || (bitsizeof(size_t) - 7) < shift) {\n \t\t\terror(\"bad object header\");\n \t\t\tsize = used = 0;\n \t\t\tbreak;\n@@ -1153,7 +1157,7 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \t\tsize = st_add(size, st_left_shift(c & 0x7f, shift));\n \t\tshift += 7;\n \t}\n-\t*sizep = cast_size_t_to_ulong(size);\n+\t*sizep = size;\n \treturn used;\n }\n \n@@ -1215,7 +1219,7 @@ unsigned long get_size_from_delta(struct packed_git *p,\n int unpack_object_header(struct packed_git *p,\n \t\t\t struct pack_window **w_curs,\n \t\t\t off_t *curpos,\n-\t\t\t unsigned long *sizep)\n+\t\t\t size_t *sizep)\n {\n \tunsigned char *base;\n \tunsigned long left;\n@@ -1367,7 +1371,7 @@ static enum object_type packed_to_object_type(struct repository *r,\n \n \twhile (type == OBJ_OFS_DELTA || type == OBJ_REF_DELTA) {\n \t\toff_t base_offset;\n-\t\tunsigned long size;\n+\t\tsize_t size;\n \t\t/* Push the object we're going to leave behind */\n \t\tif (poi_stack_nr >= poi_stack_alloc && poi_stack == small_poi_stack) {\n \t\t\tpoi_stack_alloc = alloc_nr(poi_stack_nr);\n@@ -1586,7 +1590,7 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n \t\t\t\t\t     uint32_t *maybe_index_pos, struct object_info *oi)\n {\n \tstruct pack_window *w_curs = NULL;\n-\tunsigned long size;\n+\tsize_t size;\n \toff_t curpos = obj_offset;\n \tenum object_type type = OBJ_NONE;\n \tuint32_t pack_pos;\n@@ -1778,7 +1782,7 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n \tstruct pack_window *w_curs = NULL;\n \toff_t curpos = obj_offset;\n \tvoid *data = NULL;\n-\tunsigned long size;\n+\tsize_t size;\n \tenum object_type type;\n \tstruct unpack_entry_stack_ent small_delta_stack[UNPACK_ENTRY_STACK_PREALLOC];\n \tstruct unpack_entry_stack_ent *delta_stack = small_delta_stack;\n@@ -1943,8 +1947,10 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n \t\t\t      (uintmax_t)curpos, p->pack_name);\n \t\t\tdata = NULL;\n \t\t} else {\n+\t\t\tunsigned long sz;\n \t\t\tdata = patch_delta(base, base_size, delta_data,\n-\t\t\t\t\t   delta_size, &size);\n+\t\t\t\t\t   delta_size, &sz);\n+\t\t\tsize = sz;\n \n \t\t\t/*\n \t\t\t * We could not apply the delta; warn the user, but\n@@ -2929,7 +2935,7 @@ int packfile_read_object_stream(struct odb_read_stream **out,\n \tstruct odb_packed_read_stream *stream;\n \tstruct pack_window *window = NULL;\n \tenum object_type in_pack_type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \tin_pack_type = unpack_object_header(pack, &window, &offset, &size);\n \tunuse_pack(&window);\ndiff --git a/packfile.h b/packfile.h\nindex 9b647da7dd..49d6bdecf6 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -456,9 +456,9 @@ off_t find_pack_entry_one(const struct object_id *oid, struct packed_git *);\n \n int is_pack_valid(struct packed_git *);\n void *unpack_entry(struct repository *r, struct packed_git *, off_t, enum object_type *, unsigned long *);\n-unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n+unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, size_t *sizep);\n unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n-int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\n+int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, size_t *);\n off_t get_delta_base(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t *curpos, enum object_type type,\n \t\t     off_t delta_obj_offset);\n-- \ngitgitgadget\n\n"},{"id":"542688","messageId":"3274cba862ae42a6813710410274a692ec0f5d29.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 04/11] delta, packfile: use size_t for delta header sizes","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:21Z","receivedAt":"2026-05-04T17:08:35Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe delta header decoding functions return unsigned long, which\ntruncates on Windows for objects larger than 4GB. Introduce size_t\nvariants get_delta_hdr_size_sz() and get_size_from_delta_sz() that\npreserve the full 64-bit size, and use them in packed_object_info()\nwhere the size is needed for streaming decisions.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n delta.h    | 14 ++++++++++++--\n packfile.c | 33 ++++++++++++++++++++++++---------\n 2 files changed, 36 insertions(+), 11 deletions(-)\n\ndiff --git a/delta.h b/delta.h\nindex 8a56ec0799..fad68cfc45 100644\n--- a/delta.h\n+++ b/delta.h\n@@ -86,8 +86,11 @@ void *patch_delta(const void *src_buf, unsigned long src_size,\n  * This must be called twice on the delta data buffer, first to get the\n  * expected source buffer size, and again to get the target buffer size.\n  */\n-static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n-\t\t\t\t\t       const unsigned char *top)\n+/*\n+ * Size_t variant that doesn't truncate - use for >4GB objects on Windows.\n+ */\n+static inline size_t get_delta_hdr_size_sz(const unsigned char **datap,\n+\t\t\t\t\t   const unsigned char *top)\n {\n \tconst unsigned char *data = *datap;\n \tsize_t cmd, size = 0;\n@@ -98,6 +101,13 @@ static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n \t\ti += 7;\n \t} while (cmd & 0x80 && data < top);\n \t*datap = data;\n+\treturn size;\n+}\n+\n+static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n+\t\t\t\t\t       const unsigned char *top)\n+{\n+\tsize_t size = get_delta_hdr_size_sz(datap, top);\n \treturn cast_size_t_to_ulong(size);\n }\n \ndiff --git a/packfile.c b/packfile.c\nindex fdae91dd11..4208f53046 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1161,9 +1161,12 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \treturn used;\n }\n \n-unsigned long get_size_from_delta(struct packed_git *p,\n-\t\t\t\t  struct pack_window **w_curs,\n-\t\t\t\t  off_t curpos)\n+/*\n+ * Size_t variant for >4GB delta results on Windows.\n+ */\n+static size_t get_size_from_delta_sz(struct packed_git *p,\n+\t\t\t\t     struct pack_window **w_curs,\n+\t\t\t\t     off_t curpos)\n {\n \tconst unsigned char *data;\n \tunsigned char delta_head[20], *in;\n@@ -1210,10 +1213,18 @@ unsigned long get_size_from_delta(struct packed_git *p,\n \tdata = delta_head;\n \n \t/* ignore base size */\n-\tget_delta_hdr_size(&data, delta_head+sizeof(delta_head));\n+\tget_delta_hdr_size_sz(&data, delta_head+sizeof(delta_head));\n \n \t/* Read the result size */\n-\treturn get_delta_hdr_size(&data, delta_head+sizeof(delta_head));\n+\treturn get_delta_hdr_size_sz(&data, delta_head+sizeof(delta_head));\n+}\n+\n+unsigned long get_size_from_delta(struct packed_git *p,\n+\t\t\t\t  struct pack_window **w_curs,\n+\t\t\t\t  off_t curpos)\n+{\n+\tsize_t size = get_size_from_delta_sz(p, w_curs, curpos);\n+\treturn cast_size_t_to_ulong(size);\n }\n \n int unpack_object_header(struct packed_git *p,\n@@ -1618,14 +1629,18 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n \t\t\t\tret = -1;\n \t\t\t\tgoto out;\n \t\t\t}\n-\t\t\t*oi->sizep = get_size_from_delta(p, &w_curs, tmp_pos);\n-\t\t\tif (*oi->sizep == 0) {\n+\t\t\t/*\n+\t\t\t * Use size_t variant to avoid die() on >4GB deltas.\n+\t\t\t * oi->sizep is unsigned long, so truncation may occur,\n+\t\t\t * but streaming code uses its own size_t tracking.\n+\t\t\t */\n+\t\t\tsize = get_size_from_delta_sz(p, &w_curs, tmp_pos);\n+\t\t\tif (size == 0) {\n \t\t\t\tret = -1;\n \t\t\t\tgoto out;\n \t\t\t}\n-\t\t} else {\n-\t\t\t*oi->sizep = size;\n \t\t}\n+\t\t*oi->sizep = (unsigned long)size;\n \t}\n \n \tif (oi->disk_sizep || (oi->mtimep && p->is_cruft)) {\n-- \ngitgitgadget\n\n"},{"id":"542689","messageId":"afa74a3a2b9caf9989055a9311309f590729d6c1.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 05/11] test-tool: add a helper to synthesize large packfiles","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:22Z","receivedAt":"2026-05-04T17:08:37Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nTo test Git's behavior with very large pack files, we need a way to\ngenerate such files quickly.\n\nA naive approach using only readily-available Git commands would take\nover 10 hours for a 4GB pack file, which is prohibitive.\n\nSide-stepping Git's machinery and actual zlib compression by writing\nuncompressed content with the appropriate zlib header makes things\nmuch faster. The fastest method using this approach generates many\nsmall, unreachable blob objects and takes about 1.5 minutes for 4GB.\nHowever, this cannot be used because we need to test git clone, which\nrequires a reachable commit history.\n\nGenerating many reachable commits with small, uncompressed blobs takes\nabout 4 minutes for 4GB. But this approach 1) does not reproduce the\nissues we want to fix (which require individual objects larger than\n4GB) and 2) is comparatively slow because of the many SHA-1\ncalculations.\n\nThe approach taken here generates a single large blob (filled with NUL\nbytes), along with the trees and commits needed to make it reachable.\nThis takes about 2.5 minutes for 4.5GB, which is the fastest option\nthat produces a valid, clonable repository with an object large enough\nto trigger the bugs we want to test.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Makefile                   |   1 +\n compat/zlib-compat.h       |   2 +\n t/helper/meson.build       |   1 +\n t/helper/test-synthesize.c | 250 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c       |   1 +\n t/helper/test-tool.h       |   1 +\n 6 files changed, 256 insertions(+)\n create mode 100644 t/helper/test-synthesize.c\n\ndiff --git a/Makefile b/Makefile\nindex cedc234173..85405cb5b8 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -872,6 +872,7 @@ TEST_BUILTINS_OBJS += test-submodule-config.o\n TEST_BUILTINS_OBJS += test-submodule-nested-repo-config.o\n TEST_BUILTINS_OBJS += test-submodule.o\n TEST_BUILTINS_OBJS += test-subprocess.o\n+TEST_BUILTINS_OBJS += test-synthesize.o\n TEST_BUILTINS_OBJS += test-trace2.o\n TEST_BUILTINS_OBJS += test-truncate.o\n TEST_BUILTINS_OBJS += test-userdiff.o\ndiff --git a/compat/zlib-compat.h b/compat/zlib-compat.h\nindex ac08276622..5078c5ef6c 100644\n--- a/compat/zlib-compat.h\n+++ b/compat/zlib-compat.h\n@@ -7,6 +7,8 @@\n # define z_stream_s zng_stream_s\n # define gz_header_s zng_gz_header_s\n \n+# define adler32(adler, buf, len) zng_adler32(adler, buf, len)\n+\n # define crc32(crc, buf, len) zng_crc32(crc, buf, len)\n \n # define inflate(strm, bits) zng_inflate(strm, bits)\ndiff --git a/t/helper/meson.build b/t/helper/meson.build\nindex 675e64c010..3235f10ab8 100644\n--- a/t/helper/meson.build\n+++ b/t/helper/meson.build\n@@ -69,6 +69,7 @@ test_tool_sources = [\n   'test-submodule-nested-repo-config.c',\n   'test-submodule.c',\n   'test-subprocess.c',\n+  'test-synthesize.c',\n   'test-tool.c',\n   'test-trace2.c',\n   'test-truncate.c',\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nnew file mode 100644\nindex 0000000000..3ce7078078\n--- /dev/null\n+++ b/t/helper/test-synthesize.c\n@@ -0,0 +1,250 @@\n+#define USE_THE_REPOSITORY_VARIABLE\n+\n+#include \"test-tool.h\"\n+#include \"git-compat-util.h\"\n+#include \"git-zlib.h\"\n+#include \"hash.h\"\n+#include \"hex.h\"\n+#include \"object-file.h\"\n+#include \"object.h\"\n+#include \"pack.h\"\n+#include \"parse-options.h\"\n+#include \"parse.h\"\n+#include \"repository.h\"\n+#include \"setup.h\"\n+#include \"strbuf.h\"\n+#include \"write-or-die.h\"\n+\n+#define BLOCK_SIZE 0xffff\n+static const unsigned char zeros[BLOCK_SIZE];\n+\n+/*\n+ * Write data as an uncompressed zlib stream.\n+ * For data larger than 64KB, writes multiple uncompressed blocks.\n+ * If data is NULL, writes zeros.\n+ * Updates the pack checksum context.\n+ */\n+static void write_uncompressed_zlib(FILE *f, struct git_hash_ctx *pack_ctx,\n+\t\t\t\t    const void *data, size_t len,\n+\t\t\t\t    const struct git_hash_algo *algo)\n+{\n+\tunsigned char zlib_header[2] = { 0x78, 0x01 }; /* CMF, FLG */\n+\tunsigned char block_header[5];\n+\tconst unsigned char *p = data;\n+\tsize_t remaining = len;\n+\tuint32_t adler = 1L; /* adler32 initial value */\n+\tunsigned char adler_buf[4];\n+\n+\t/* Write zlib header */\n+\tfwrite_or_die(f, zlib_header, sizeof(zlib_header));\n+\talgo->update_fn(pack_ctx, zlib_header, 2);\n+\n+\t/* Write uncompressed blocks (max 64KB each) */\n+\tdo {\n+\t\tsize_t block_len = remaining > BLOCK_SIZE ? BLOCK_SIZE : remaining;\n+\t\tint is_final = (block_len == remaining);\n+\t\tconst unsigned char *block_data = data ? p : zeros;\n+\n+\t\tblock_header[0] = is_final ? 0x01 : 0x00;\n+\t\tblock_header[1] = block_len & 0xff;\n+\t\tblock_header[2] = (block_len >> 8) & 0xff;\n+\t\tblock_header[3] = block_header[1] ^ 0xff;\n+\t\tblock_header[4] = block_header[2] ^ 0xff;\n+\n+\t\tfwrite_or_die(f, block_header, sizeof(block_header));\n+\t\talgo->update_fn(pack_ctx, block_header, 5);\n+\n+\t\tif (block_len) {\n+\t\t\tfwrite_or_die(f, block_data, block_len);\n+\t\t\talgo->update_fn(pack_ctx, block_data, block_len);\n+\t\t\tadler = adler32(adler, block_data, block_len);\n+\t\t}\n+\n+\t\tif (data)\n+\t\t\tp += block_len;\n+\t\tremaining -= block_len;\n+\t} while (remaining > 0);\n+\n+\t/* Write adler32 checksum */\n+\tput_be32(adler_buf, adler);\n+\tfwrite_or_die(f, adler_buf, sizeof(adler_buf));\n+\talgo->update_fn(pack_ctx, adler_buf, 4);\n+}\n+\n+/*\n+ * Write an uncompressed object to the pack file.\n+ * If `data == NULL`, it is treated like a buffer to NUL bytes.\n+ * Updates the pack checksum context.\n+ */\n+static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n+\t\t\t      enum object_type type,\n+\t\t\t      const void *data, size_t len,\n+\t\t\t      struct object_id *oid,\n+\t\t\t      const struct git_hash_algo *algo)\n+{\n+\tunsigned char pack_header[MAX_PACK_OBJECT_HEADER];\n+\tchar object_header[32];\n+\tint pack_header_len, object_header_len;\n+\tstruct git_hash_ctx ctx;\n+\n+\t/* Write pack object header */\n+\tpack_header_len = encode_in_pack_object_header(pack_header,\n+\t\t\t\t\t\t       sizeof(pack_header),\n+\t\t\t\t\t\t       type, len);\n+\tfwrite_or_die(f, pack_header, pack_header_len);\n+\talgo->update_fn(pack_ctx, pack_header, pack_header_len);\n+\n+\t/* Write the data as uncompressed zlib */\n+\twrite_uncompressed_zlib(f, pack_ctx, data, len, algo);\n+\n+\talgo->init_fn(&ctx);\n+\tobject_header_len = format_object_header(object_header,\n+\t\t\t\t\t\t sizeof(object_header),\n+\t\t\t\t\t\t type, len);\n+\talgo->update_fn(&ctx, object_header, object_header_len);\n+\tif (data)\n+\t\talgo->update_fn(&ctx, data, len);\n+\telse {\n+\t\tfor (size_t i = len / BLOCK_SIZE; i; i--)\n+\t\t\talgo->update_fn(&ctx, zeros, BLOCK_SIZE);\n+\t\talgo->update_fn(&ctx, zeros, len % BLOCK_SIZE);\n+\t}\n+\talgo->final_oid_fn(oid, &ctx);\n+}\n+\n+/*\n+ * Generate a pack file with a single large (>4GB) reachable object.\n+ *\n+ * Creates:\n+ *   1. A large blob (all NUL bytes)\n+ *   2. A tree containing that blob as \"file\"\n+ *   3. A commit using that tree\n+ *   4. The empty tree\n+ *   5. A child commit using the empty tree\n+ *\n+ * This is useful for testing that Git can handle objects larger than 4GB.\n+ */\n+static int generate_pack_with_large_object(const char *path, size_t blob_size,\n+\t\t\t\t\t   const struct git_hash_algo *algo)\n+{\n+\tFILE *f = xfopen(path, \"wb\");\n+\tstruct git_hash_ctx pack_ctx;\n+\tunsigned char pack_hash[GIT_MAX_RAWSZ];\n+\tstruct object_id blob_oid, tree_oid, commit_oid, empty_tree_oid, final_commit_oid;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tconst uint32_t object_count = 5;\n+\tstruct pack_header pack_header = {\n+\t\t.hdr_signature = htonl(PACK_SIGNATURE),\n+\t\t.hdr_version = htonl(PACK_VERSION),\n+\t\t.hdr_entries = htonl(object_count),\n+\t};\n+\n+\talgo->init_fn(&pack_ctx);\n+\n+\t/* Write pack header */\n+\tfwrite_or_die(f, &pack_header, sizeof(pack_header));\n+\talgo->update_fn(&pack_ctx, &pack_header, sizeof(pack_header));\n+\n+\t/* 1. Write the large blob */\n+\twrite_pack_object(f, &pack_ctx, OBJ_BLOB, NULL, blob_size, &blob_oid, algo);\n+\n+\t/* 2. Write tree containing the blob as \"file\" */\n+\tstrbuf_addf(&buf, \"100644 file%c\", '\\0');\n+\tstrbuf_add(&buf, blob_oid.hash, algo->rawsz);\n+\twrite_pack_object(f, &pack_ctx, OBJ_TREE, buf.buf, buf.len, &tree_oid, algo);\n+\n+\t/* 3. Write commit using that tree */\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf,\n+\t\t    \"tree %s\\n\"\n+\t\t    \"author A U Thor <author@example.com> 1234567890 +0000\\n\"\n+\t\t    \"committer C O Mitter <committer@example.com> 1234567890 +0000\\n\"\n+\t\t    \"\\n\"\n+\t\t    \"Large blob commit\\n\",\n+\t\t    oid_to_hex(&tree_oid));\n+\twrite_pack_object(f, &pack_ctx, OBJ_COMMIT, buf.buf, buf.len, &commit_oid, algo);\n+\n+\t/* 4. Write the empty tree */\n+\twrite_pack_object(f, &pack_ctx, OBJ_TREE, \"\", 0, &empty_tree_oid, algo);\n+\n+\t/* 5. Write final commit using empty tree, with previous commit as parent */\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf,\n+\t\t    \"tree %s\\n\"\n+\t\t    \"parent %s\\n\"\n+\t\t    \"author A U Thor <author@example.com> 1234567890 +0000\\n\"\n+\t\t    \"committer C O Mitter <committer@example.com> 1234567890 +0000\\n\"\n+\t\t    \"\\n\"\n+\t\t    \"Empty tree commit\\n\",\n+\t\t    oid_to_hex(&empty_tree_oid),\n+\t\t    oid_to_hex(&commit_oid));\n+\twrite_pack_object(f, &pack_ctx, OBJ_COMMIT, buf.buf, buf.len, &final_commit_oid, algo);\n+\n+\t/* Write pack trailer (checksum) */\n+\talgo->final_fn(pack_hash, &pack_ctx);\n+\tfwrite_or_die(f, pack_hash, algo->rawsz);\n+\tif (fclose(f))\n+\t\tdie_errno(_(\"could not close '%s'\"), path);\n+\n+\tstrbuf_release(&buf);\n+\n+\t/* Print the final commit OID so caller can set up refs */\n+\tprintf(\"%s\\n\", oid_to_hex(&final_commit_oid));\n+\n+\treturn 0;\n+}\n+\n+static int cmd__synthesize__pack(int argc, const char **argv,\n+\t\t\t\t const char *prefix UNUSED,\n+\t\t\t\t struct repository *repo)\n+{\n+\tint non_git;\n+\tint reachable_large = 0;\n+\tconst struct git_hash_algo *algo;\n+\tsize_t blob_size;\n+\tuintmax_t blob_size_u;\n+\tconst char *path;\n+\tconst char * const usage[] = {\n+\t\t\"test-tool synthesize pack \"\n+\t\t\"--reachable-large <blob-size> <filename>\",\n+\t\tNULL\n+\t};\n+\tstruct option options[] = {\n+\t\tOPT_BOOL(0, \"reachable-large\", &reachable_large,\n+\t\t\t N_(\"write a pack with a single reachable large blob\")),\n+\t\tOPT_END()\n+\t};\n+\n+\tsetup_git_directory_gently(&non_git);\n+\trepo = the_repository;\n+\talgo = repo->hash_algo;\n+\n+\targc = parse_options(argc, argv, NULL, options, usage,\n+\t\t\t     PARSE_OPT_KEEP_ARGV0);\n+\tif (argc != 3 || !reachable_large)\n+\t\tusage_with_options(usage, options);\n+\n+\tif (!git_parse_unsigned(argv[1], &blob_size_u,\n+\t\t\t\tmaximum_unsigned_value_of_type(size_t)))\n+\t\tdie(_(\"'%s' is not a valid blob size\"), argv[1]);\n+\tblob_size = blob_size_u;\n+\tpath = argv[2];\n+\n+\treturn !!generate_pack_with_large_object(path, blob_size, algo);\n+}\n+\n+int cmd__synthesize(int argc, const char **argv)\n+{\n+\tconst char *prefix = NULL;\n+\tchar const * const synthesize_usage[] = {\n+\t\t\"test-tool synthesize pack <options>\",\n+\t\tNULL,\n+\t};\n+\tparse_opt_subcommand_fn *fn = NULL;\n+\tstruct option options[] = {\n+\t\tOPT_SUBCOMMAND(\"pack\", &fn, cmd__synthesize__pack),\n+\t\tOPT_END()\n+\t};\n+\targc = parse_options(argc, argv, prefix, options, synthesize_usage, 0);\n+\treturn !!fn(argc, argv, prefix, NULL);\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex a7abc618b3..b71a22b43b 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -82,6 +82,7 @@ static struct test_cmd cmds[] = {\n \t{ \"submodule-config\", cmd__submodule_config },\n \t{ \"submodule-nested-repo-config\", cmd__submodule_nested_repo_config },\n \t{ \"subprocess\", cmd__subprocess },\n+\t{ \"synthesize\", cmd__synthesize },\n \t{ \"trace2\", cmd__trace2 },\n \t{ \"truncate\", cmd__truncate },\n \t{ \"userdiff\", cmd__userdiff },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 7f150fa1eb..f2885b33d5 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -75,6 +75,7 @@ int cmd__submodule(int argc, const char **argv);\n int cmd__submodule_config(int argc, const char **argv);\n int cmd__submodule_nested_repo_config(int argc, const char **argv);\n int cmd__subprocess(int argc, const char **argv);\n+int cmd__synthesize(int argc, const char **argv);\n int cmd__trace2(int argc, const char **argv);\n int cmd__truncate(int argc, const char **argv);\n int cmd__userdiff(int argc, const char **argv);\n-- \ngitgitgadget\n\n"},{"id":"542690","messageId":"a3019888d8465e0f77926a91a20db170fef6989d.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 06/11] t5608: add regression test for >4GB object clone","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:23Z","receivedAt":"2026-05-04T17:08:38Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe shift overflow bug in index-pack and unpack-objects caused incorrect\nobject size calculation when the encoded size required more than 32 bits\nof shift. This would result in corrupted or failed unpacking of objects\nlarger than 4GB.\n\nAdd a test that creates a pack file containing a 4GB+ blob using the\nnew 'test-tool synthesize pack --reachable-large' command, then clones\nthe repository to verify the fix works correctly.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t5608-clone-2gb.sh | 37 +++++++++++++++++++++++++++++++++++++\n 1 file changed, 37 insertions(+)\n\ndiff --git a/t/t5608-clone-2gb.sh b/t/t5608-clone-2gb.sh\nindex 87a8cd9f98..af93302dde 100755\n--- a/t/t5608-clone-2gb.sh\n+++ b/t/t5608-clone-2gb.sh\n@@ -49,4 +49,41 @@ test_expect_success 'clone - with worktree, file:// protocol' '\n \n '\n \n+test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n+\tlarge_blob_size=$((4*1024*1024*1024+1)) &&\n+\tgit init --bare 4gb-repo &&\n+\thead_oid=$(test-tool synthesize pack \\\n+\t\t--reachable-large \"$large_blob_size\" \\\n+\t\t4gb-repo/objects/pack/test.pack) &&\n+\tgit -C 4gb-repo index-pack objects/pack/test.pack &&\n+\tgit -C 4gb-repo update-ref refs/heads/main $head_oid &&\n+\tgit -C 4gb-repo symbolic-ref HEAD refs/heads/main\n+'\n+\n+test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n+\t# The synthesized pack has five objects, so a large unpack limit keeps\n+\t# fetch-pack on the unpack-objects path.\n+\tgit -c fetch.unpackLimit=100 clone --bare \\\n+\t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-unpack &&\n+\n+\t# Verify the large blob survived the clone by comparing its OID\n+\t# between source and clone.  We cannot use \"cat-file -s\" because\n+\t# object_info.sizep is still unsigned long, which truncates >4GB\n+\t# sizes on Windows.  OID equality proves content integrity since\n+\t# the clone already verified checksums via index-pack/unpack-objects.\n+\tsource_blob=$(git -C 4gb-repo rev-parse main^:file) &&\n+\tclone_blob=$(git -C 4gb-clone-unpack rev-parse main^:file) &&\n+\ttest \"$source_blob\" = \"$clone_blob\"\n+'\n+\n+test_expect_success SIZE_T_IS_64BIT 'clone with >4GB object via index-pack' '\n+\t# Force fetch-pack to hand the pack to index-pack instead.\n+\tgit -c fetch.unpackLimit=1 clone --bare \\\n+\t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-index &&\n+\n+\tsource_blob=$(git -C 4gb-repo rev-parse main^:file) &&\n+\tclone_blob=$(git -C 4gb-clone-index rev-parse main^:file) &&\n+\ttest \"$source_blob\" = \"$clone_blob\"\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"542691","messageId":"859e93e7a9f1d5ba965dc2b9891b5885a6c167ef.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 07/11] test-tool synthesize: use the unsafe hash for speed","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:24Z","receivedAt":"2026-05-04T17:08:40Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nJeff King pointed out on the mailing list [1] that t5608's new >4GB\ntest cases dominate the entire test suite runtime: 160 seconds on his\nlaptop when the rest of the suite finishes in under 90 seconds, and\n305-850 seconds across CI jobs. The bottleneck is that the synthesize\nhelper hashes roughly 8 GB of data through SHA-1 (4 GB for the pack\nchecksum plus 4 GB for the blob OID) for a 4 GB+1 blob.\n\nSince the helper generates known test data, collision detection is\nunnecessary. Switch from repo->hash_algo to unsafe_hash_algo(), which\nuses hardware-accelerated SHA-1 (via OpenSSL or Apple CommonCrypto)\nwhen available.\n\nBenchmarks on an x86_64 machine generating a 4 GB+1 pack (2 runs\neach, interleaved):\n\n  SHA-1 backend      Run 1    Run 2\n  SHA1DC (safe)       75s      80s\n  OpenSSL (unsafe)    21s      19s\n\nThe effect scales linearly. At 64 MB with 10 randomized interleaved\nruns, the OpenSSL unsafe backend shows a 5.4x improvement (median\n0.202s vs 1.088s) with tight variance (stdev 0.028s vs 0.095s).\n\nThe speedup is only realized when the build has a fast unsafe backend\ncompiled in. The CI's linux-TEST-vars job already sets\nOPENSSL_SHA1_UNSAFE=YesPlease; macOS benefits from Apple CommonCrypto\nwhen configured. On builds without a separate unsafe backend (such as\nthe default Windows builds), unsafe_hash_algo() returns the regular\ncollision-detecting implementation and the change is a no-op.\n\n[1] https://lore.kernel.org/git/20260501063805.GA2038915@coredump.intra.peff.net/\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/helper/test-synthesize.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nindex 3ce7078078..e2faaad7b4 100644\n--- a/t/helper/test-synthesize.c\n+++ b/t/helper/test-synthesize.c\n@@ -217,7 +217,7 @@ static int cmd__synthesize__pack(int argc, const char **argv,\n \n \tsetup_git_directory_gently(&non_git);\n \trepo = the_repository;\n-\talgo = repo->hash_algo;\n+\talgo = unsafe_hash_algo(repo->hash_algo);\n \n \targc = parse_options(argc, argv, NULL, options, usage,\n \t\t\t     PARSE_OPT_KEEP_ARGV0);\n-- \ngitgitgadget\n\n"},{"id":"542692","messageId":"29b9a74e915e6200ac2b4d98e446c1e73964cbd2.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 08/11] test-tool synthesize: precompute pack for 4 GiB + 1","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:25Z","receivedAt":"2026-05-04T17:08:40Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe synthesize helper hashes roughly 8 GiB of data through SHA-1 to\nproduce a 4 GiB + 1 pack (4 GiB for the pack checksum, 4 GiB for\nthe blob OID). Since the blob content is all NUL bytes, every byte\nin the resulting pack file is deterministic for a given blob size and\nhash algorithm.\n\nAdd a fast path that writes the pack from precomputed constants:\na 25-byte prefix (pack header, object header, zlib header, first\nblock header), the zero-filled bulk with periodic 5-byte deflate\nblock headers, and a 513-byte suffix (tree, two commits, empty tree,\npack SHA-1 checksum). This eliminates all SHA-1 and adler32\ncomputation, making the helper purely I/O-bound.\n\nThe precomputed constants are stored in a struct fast_pack array\nkeyed by hash algorithm format_id, so that adding SHA-256 support\nlater requires only adding another array entry with its suffix.\n\nThe constants were generated by running the generic path and\nextracting the non-zero bytes from the resulting pack file.\n\nBenchmarks generating a 4 GiB + 1 pack (3 runs each, SHA1DC on\nx86_64):\n\n  generic path:   88s / 81s / 140s\n  fast path:      14s / 13s / 15s\n\nOn CI, where t5608 currently takes 200-850 seconds depending on the\njob, the fast path cuts the pack-generation phase from minutes to\nseconds, leaving only the clone operations themselves.\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/helper/test-synthesize.c | 202 ++++++++++++++++++++++++++++++++++++-\n 1 file changed, 201 insertions(+), 1 deletion(-)\n\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nindex e2faaad7b4..83c40ee02a 100644\n--- a/t/helper/test-synthesize.c\n+++ b/t/helper/test-synthesize.c\n@@ -112,6 +112,201 @@ static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n \talgo->final_oid_fn(oid, &ctx);\n }\n \n+/*\n+ * Fast path: precomputed pack data for a 4 GiB + 1 all-NUL blob.\n+ *\n+ * The generated pack is almost entirely zeros with a small constant\n+ * prefix, periodic deflate block headers, and a constant suffix\n+ * containing the tree, two commits, and the pack checksum.  Because\n+ * every byte is deterministic for a given blob size and hash algorithm,\n+ * we can write the pack without computing any hashes at all, reducing\n+ * runtime from minutes of hash computation to seconds of pure I/O.\n+ *\n+ * The blob is stored as an uncompressed deflate stream: a two-byte\n+ * zlib header, then 65538 blocks of up to 0xffff bytes each, followed\n+ * by an adler32 checksum.  The pack header and deflate framing are\n+ * shared across hash algorithms; only the suffix (which contains OIDs\n+ * and the pack checksum) differs.\n+ *\n+ * Constants were generated by running the generic path and extracting\n+ * the non-zero bytes from the resulting pack file.\n+ */\n+\n+#define FAST_PACK_4G1_BLOB_SIZE ((size_t)4 * 1024 * 1024 * 1024 + 1)\n+#define FAST_PACK_4G1_N_FULL_BLOCKS 65537\n+\n+/*\n+ * Per-hash-algorithm constants for the fast path.  The prefix and\n+ * deflate block structure are identical across algorithms; only the\n+ * suffix (tree, commits, pack checksum) and the commit OID differ.\n+ */\n+struct fast_pack {\n+\tuint32_t format_id;\n+\tconst unsigned char *suffix;\n+\tsize_t suffix_len;\n+\tconst char *commit_oid;\n+};\n+\n+/* Pack header + pack object header + zlib header + first block header */\n+static const unsigned char fast_pack_prefix[] = {\n+\t/* PACK header: signature, version 2, 5 objects */\n+\t0x50, 0x41, 0x43, 0x4b, 0x00, 0x00, 0x00, 0x02,\n+\t0x00, 0x00, 0x00, 0x05,\n+\t/* pack object header: blob, size = 4294967297 */\n+\t0xb1, 0x80, 0x80, 0x80, 0x80, 0x01,\n+\t/* zlib header: CMF=0x78, FLG=0x01 */\n+\t0x78, 0x01,\n+\t/* first non-final block header: BFINAL=0, LEN=0xffff, NLEN=0x0000 */\n+\t0x00, 0xff, 0xff, 0x00, 0x00\n+};\n+\n+/* Every non-final deflate block header is identical */\n+static const unsigned char fast_pack_block_header[] = {\n+\t0x00, 0xff, 0xff, 0x00, 0x00\n+};\n+\n+/* Final block (2 data bytes) + adler32 of 4294967297 NUL bytes */\n+static const unsigned char fast_pack_final_block[] = {\n+\t/* BFINAL=1, LEN=2, NLEN=0xfffd */\n+\t0x01, 0x02, 0x00, 0xfd, 0xff,\n+\t/* 2 NUL data bytes */\n+\t0x00, 0x00,\n+\t/* adler32 */\n+\t0x00, 0xe2, 0x00, 0x01\n+};\n+\n+/*\n+ * SHA-1 suffix: tree, commit, empty tree, final commit, pack checksum.\n+ */\n+static const unsigned char fast_pack_sha1_suffix[] = {\n+\t0xa0, 0x02, 0x78, 0x01, 0x01, 0x20, 0x00, 0xdf,\n+\t0xff, 0x31, 0x30, 0x30, 0x36, 0x34, 0x34, 0x20,\n+\t0x66, 0x69, 0x6c, 0x65, 0x00, 0x3e, 0xb7, 0xfe,\n+\t0xb1, 0x41, 0x3c, 0x75, 0x7f, 0x0d, 0x81, 0x81,\n+\t0xde, 0xb2, 0x8d, 0x1d, 0xab, 0x03, 0xd6, 0x48,\n+\t0x46, 0xb4, 0xb4, 0x0c, 0x60, 0x95, 0x0b, 0x78,\n+\t0x01, 0x01, 0xb5, 0x00, 0x4a, 0xff, 0x74, 0x72,\n+\t0x65, 0x65, 0x20, 0x63, 0x36, 0x38, 0x33, 0x66,\n+\t0x63, 0x63, 0x37, 0x64, 0x31, 0x64, 0x38, 0x33,\n+\t0x65, 0x66, 0x32, 0x66, 0x65, 0x31, 0x61, 0x66,\n+\t0x35, 0x35, 0x32, 0x31, 0x35, 0x64, 0x30, 0x31,\n+\t0x36, 0x38, 0x64, 0x62, 0x35, 0x32, 0x61, 0x33,\n+\t0x61, 0x33, 0x62, 0x0a, 0x61, 0x75, 0x74, 0x68,\n+\t0x6f, 0x72, 0x20, 0x41, 0x20, 0x55, 0x20, 0x54,\n+\t0x68, 0x6f, 0x72, 0x20, 0x3c, 0x61, 0x75, 0x74,\n+\t0x68, 0x6f, 0x72, 0x40, 0x65, 0x78, 0x61, 0x6d,\n+\t0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e,\n+\t0x20, 0x31, 0x32, 0x33, 0x34, 0x35, 0x36, 0x37,\n+\t0x38, 0x39, 0x30, 0x20, 0x2b, 0x30, 0x30, 0x30,\n+\t0x30, 0x0a, 0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74,\n+\t0x74, 0x65, 0x72, 0x20, 0x43, 0x20, 0x4f, 0x20,\n+\t0x4d, 0x69, 0x74, 0x74, 0x65, 0x72, 0x20, 0x3c,\n+\t0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65,\n+\t0x72, 0x40, 0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c,\n+\t0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31,\n+\t0x32, 0x33, 0x34, 0x35, 0x36, 0x37, 0x38, 0x39,\n+\t0x30, 0x20, 0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a,\n+\t0x0a, 0x4c, 0x61, 0x72, 0x67, 0x65, 0x20, 0x62,\n+\t0x6c, 0x6f, 0x62, 0x20, 0x63, 0x6f, 0x6d, 0x6d,\n+\t0x69, 0x74, 0x0a, 0xc6, 0x55, 0x37, 0x6b, 0x20,\n+\t0x78, 0x01, 0x01, 0x00, 0x00, 0xff, 0xff, 0x00,\n+\t0x00, 0x00, 0x01, 0x95, 0x0e, 0x78, 0x01, 0x01,\n+\t0xe5, 0x00, 0x1a, 0xff, 0x74, 0x72, 0x65, 0x65,\n+\t0x20, 0x34, 0x62, 0x38, 0x32, 0x35, 0x64, 0x63,\n+\t0x36, 0x34, 0x32, 0x63, 0x62, 0x36, 0x65, 0x62,\n+\t0x39, 0x61, 0x30, 0x36, 0x30, 0x65, 0x35, 0x34,\n+\t0x62, 0x66, 0x38, 0x64, 0x36, 0x39, 0x32, 0x38,\n+\t0x38, 0x66, 0x62, 0x65, 0x65, 0x34, 0x39, 0x30,\n+\t0x34, 0x0a, 0x70, 0x61, 0x72, 0x65, 0x6e, 0x74,\n+\t0x20, 0x63, 0x35, 0x62, 0x32, 0x31, 0x63, 0x36,\n+\t0x31, 0x31, 0x61, 0x61, 0x35, 0x39, 0x34, 0x65,\n+\t0x63, 0x39, 0x66, 0x64, 0x37, 0x65, 0x39, 0x32,\n+\t0x63, 0x66, 0x39, 0x36, 0x34, 0x38, 0x39, 0x31,\n+\t0x34, 0x63, 0x61, 0x34, 0x63, 0x32, 0x34, 0x31,\n+\t0x32, 0x0a, 0x61, 0x75, 0x74, 0x68, 0x6f, 0x72,\n+\t0x20, 0x41, 0x20, 0x55, 0x20, 0x54, 0x68, 0x6f,\n+\t0x72, 0x20, 0x3c, 0x61, 0x75, 0x74, 0x68, 0x6f,\n+\t0x72, 0x40, 0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c,\n+\t0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31,\n+\t0x32, 0x33, 0x34, 0x35, 0x36, 0x37, 0x38, 0x39,\n+\t0x30, 0x20, 0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a,\n+\t0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65,\n+\t0x72, 0x20, 0x43, 0x20, 0x4f, 0x20, 0x4d, 0x69,\n+\t0x74, 0x74, 0x65, 0x72, 0x20, 0x3c, 0x63, 0x6f,\n+\t0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65, 0x72, 0x40,\n+\t0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e,\n+\t0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33,\n+\t0x34, 0x35, 0x36, 0x37, 0x38, 0x39, 0x30, 0x20,\n+\t0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a, 0x0a, 0x45,\n+\t0x6d, 0x70, 0x74, 0x79, 0x20, 0x74, 0x72, 0x65,\n+\t0x65, 0x20, 0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74,\n+\t0x0a, 0xaa, 0xb8, 0x45, 0x01, 0x8e, 0xfc, 0xf0,\n+\t0x2f, 0x9c, 0xc5, 0xcc, 0x4f, 0x6a, 0x1a, 0xc9,\n+\t0x2b, 0x23, 0xa9, 0xff, 0x91, 0x06, 0xc2, 0x70,\n+\t0xe3\n+};\n+\n+static const struct fast_pack fast_packs[] = {\n+\t{\n+\t\t.format_id = GIT_SHA1_FORMAT_ID,\n+\t\t.suffix = fast_pack_sha1_suffix,\n+\t\t.suffix_len = sizeof(fast_pack_sha1_suffix),\n+\t\t.commit_oid = \"aac43daf40d0377af31aa9c798a4ae8a31b55c1d\",\n+\t},\n+};\n+\n+/*\n+ * Try the fast path for known blob sizes.  Returns 1 if the pack was\n+ * written from precomputed constants, 0 if the caller should fall\n+ * through to the generic path.\n+ */\n+static int generate_fast_pack(const char *path, size_t blob_size,\n+\t\t\t      const struct git_hash_algo *algo)\n+{\n+\tconst struct fast_pack *fp = NULL;\n+\tFILE *f;\n+\tsize_t i;\n+\n+\tif (blob_size != FAST_PACK_4G1_BLOB_SIZE)\n+\t\treturn 0;\n+\n+\tfor (i = 0; i < ARRAY_SIZE(fast_packs); i++) {\n+\t\tif (fast_packs[i].format_id == algo->format_id) {\n+\t\t\tfp = &fast_packs[i];\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\tif (!fp)\n+\t\treturn 0;\n+\n+\tf = xfopen(path, \"wb\");\n+\n+\tfwrite_or_die(f, fast_pack_prefix, sizeof(fast_pack_prefix));\n+\n+\t/* First full block: 0xffff zero bytes (header already in prefix) */\n+\tfwrite_or_die(f, zeros, BLOCK_SIZE);\n+\n+\t/* Remaining non-final full blocks */\n+\tfor (i = 1; i < FAST_PACK_4G1_N_FULL_BLOCKS; i++) {\n+\t\tfwrite_or_die(f, fast_pack_block_header,\n+\t\t\t      sizeof(fast_pack_block_header));\n+\t\tfwrite_or_die(f, zeros, BLOCK_SIZE);\n+\t}\n+\n+\t/* Final block (2 data bytes) + adler32 */\n+\tfwrite_or_die(f, fast_pack_final_block,\n+\t\t      sizeof(fast_pack_final_block));\n+\n+\t/* Tree, commits, and pack checksum */\n+\tfwrite_or_die(f, fp->suffix, fp->suffix_len);\n+\n+\tif (fclose(f))\n+\t\tdie_errno(_(\"could not close '%s'\"), path);\n+\n+\tprintf(\"%s\\n\", fp->commit_oid);\n+\treturn 1;\n+}\n+\n /*\n  * Generate a pack file with a single large (>4GB) reachable object.\n  *\n@@ -127,7 +322,7 @@ static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n static int generate_pack_with_large_object(const char *path, size_t blob_size,\n \t\t\t\t\t   const struct git_hash_algo *algo)\n {\n-\tFILE *f = xfopen(path, \"wb\");\n+\tFILE *f;\n \tstruct git_hash_ctx pack_ctx;\n \tunsigned char pack_hash[GIT_MAX_RAWSZ];\n \tstruct object_id blob_oid, tree_oid, commit_oid, empty_tree_oid, final_commit_oid;\n@@ -139,6 +334,11 @@ static int generate_pack_with_large_object(const char *path, size_t blob_size,\n \t\t.hdr_entries = htonl(object_count),\n \t};\n \n+\tif (generate_fast_pack(path, blob_size, algo))\n+\t\treturn 0;\n+\n+\tf = xfopen(path, \"wb\");\n+\n \talgo->init_fn(&pack_ctx);\n \n \t/* Write pack header */\n-- \ngitgitgadget\n\n"},{"id":"542693","messageId":"8e6e7208040917a254379fd6c63d432f5e2f6f59.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 09/11] test-tool synthesize: add precomputed SHA-256 pack for 4 GiB + 1","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:26Z","receivedAt":"2026-05-04T17:08:42Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nAdd a SHA-256 entry to the fast_packs[] table. The pack prefix and\ndeflate block structure are identical to SHA-1 (the pack format does\nnot encode the hash algorithm in its header). Only the suffix differs:\nSHA-256 OIDs are 32 bytes instead of 20, giving a 609-byte suffix\ncompared to 513 for SHA-1, and a different pack checksum.\n\nThe constants were generated by running the generic path inside a\nrepository initialized with --object-format=sha256.\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/helper/test-synthesize.c | 91 ++++++++++++++++++++++++++++++++++++++\n 1 file changed, 91 insertions(+)\n\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nindex 83c40ee02a..1f28ecf0f2 100644\n--- a/t/helper/test-synthesize.c\n+++ b/t/helper/test-synthesize.c\n@@ -246,6 +246,90 @@ static const unsigned char fast_pack_sha1_suffix[] = {\n \t0xe3\n };\n \n+/*\n+ * SHA-256 suffix: same structure, but with 32-byte OIDs and SHA-256\n+ * pack checksum (609 bytes vs 513 for SHA-1).\n+ */\n+static const unsigned char fast_pack_sha256_suffix[] = {\n+\t0xac, 0x02, 0x78, 0x01, 0x01, 0x2c, 0x00, 0xd3,\n+\t0xff, 0x31, 0x30, 0x30, 0x36, 0x34, 0x34, 0x20,\n+\t0x66, 0x69, 0x6c, 0x65, 0x00, 0x42, 0x53, 0xc1,\n+\t0x8a, 0x9f, 0x5e, 0xc3, 0xbb, 0x47, 0xb0, 0x83,\n+\t0x8a, 0x19, 0xdb, 0x31, 0xbb, 0x7b, 0x0f, 0x3b,\n+\t0x80, 0xa4, 0xbc, 0x2f, 0xaf, 0x72, 0x6b, 0xdb,\n+\t0x62, 0xaa, 0xba, 0xdd, 0xde, 0x77, 0xc6, 0x13,\n+\t0xeb, 0x9d, 0x0c, 0x78, 0x01, 0x01, 0xcd, 0x00,\n+\t0x32, 0xff, 0x74, 0x72, 0x65, 0x65, 0x20, 0x62,\n+\t0x36, 0x30, 0x39, 0x37, 0x37, 0x64, 0x37, 0x63,\n+\t0x34, 0x63, 0x32, 0x64, 0x31, 0x65, 0x63, 0x63,\n+\t0x33, 0x66, 0x62, 0x61, 0x31, 0x64, 0x39, 0x38,\n+\t0x65, 0x65, 0x31, 0x32, 0x30, 0x61, 0x64, 0x63,\n+\t0x32, 0x34, 0x38, 0x33, 0x34, 0x39, 0x35, 0x30,\n+\t0x62, 0x65, 0x34, 0x31, 0x32, 0x64, 0x39, 0x34,\n+\t0x63, 0x38, 0x30, 0x39, 0x34, 0x38, 0x30, 0x66,\n+\t0x35, 0x38, 0x62, 0x61, 0x39, 0x64, 0x61, 0x0a,\n+\t0x61, 0x75, 0x74, 0x68, 0x6f, 0x72, 0x20, 0x41,\n+\t0x20, 0x55, 0x20, 0x54, 0x68, 0x6f, 0x72, 0x20,\n+\t0x3c, 0x61, 0x75, 0x74, 0x68, 0x6f, 0x72, 0x40,\n+\t0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e,\n+\t0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33,\n+\t0x34, 0x35, 0x36, 0x37, 0x38, 0x39, 0x30, 0x20,\n+\t0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a, 0x63, 0x6f,\n+\t0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65, 0x72, 0x20,\n+\t0x43, 0x20, 0x4f, 0x20, 0x4d, 0x69, 0x74, 0x74,\n+\t0x65, 0x72, 0x20, 0x3c, 0x63, 0x6f, 0x6d, 0x6d,\n+\t0x69, 0x74, 0x74, 0x65, 0x72, 0x40, 0x65, 0x78,\n+\t0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f,\n+\t0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33, 0x34, 0x35,\n+\t0x36, 0x37, 0x38, 0x39, 0x30, 0x20, 0x2b, 0x30,\n+\t0x30, 0x30, 0x30, 0x0a, 0x0a, 0x4c, 0x61, 0x72,\n+\t0x67, 0x65, 0x20, 0x62, 0x6c, 0x6f, 0x62, 0x20,\n+\t0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74, 0x0a, 0xb7,\n+\t0x80, 0x3d, 0xd7, 0x20, 0x78, 0x01, 0x01, 0x00,\n+\t0x00, 0xff, 0xff, 0x00, 0x00, 0x00, 0x01, 0x95,\n+\t0x11, 0x78, 0x01, 0x01, 0x15, 0x01, 0xea, 0xfe,\n+\t0x74, 0x72, 0x65, 0x65, 0x20, 0x36, 0x65, 0x66,\n+\t0x31, 0x39, 0x62, 0x34, 0x31, 0x32, 0x32, 0x35,\n+\t0x63, 0x35, 0x33, 0x36, 0x39, 0x66, 0x31, 0x63,\n+\t0x31, 0x30, 0x34, 0x64, 0x34, 0x35, 0x64, 0x38,\n+\t0x64, 0x38, 0x35, 0x65, 0x66, 0x61, 0x39, 0x62,\n+\t0x30, 0x35, 0x37, 0x62, 0x35, 0x33, 0x62, 0x31,\n+\t0x34, 0x62, 0x34, 0x62, 0x39, 0x62, 0x39, 0x33,\n+\t0x39, 0x64, 0x64, 0x37, 0x34, 0x64, 0x65, 0x63,\n+\t0x63, 0x35, 0x33, 0x32, 0x31, 0x0a, 0x70, 0x61,\n+\t0x72, 0x65, 0x6e, 0x74, 0x20, 0x37, 0x35, 0x62,\n+\t0x66, 0x30, 0x63, 0x34, 0x37, 0x61, 0x65, 0x34,\n+\t0x62, 0x62, 0x33, 0x30, 0x38, 0x65, 0x37, 0x63,\n+\t0x63, 0x32, 0x34, 0x38, 0x32, 0x65, 0x32, 0x32,\n+\t0x65, 0x66, 0x61, 0x65, 0x33, 0x37, 0x38, 0x37,\n+\t0x61, 0x39, 0x36, 0x38, 0x34, 0x38, 0x62, 0x64,\n+\t0x31, 0x37, 0x34, 0x39, 0x35, 0x36, 0x37, 0x31,\n+\t0x34, 0x37, 0x31, 0x35, 0x32, 0x34, 0x36, 0x64,\n+\t0x64, 0x62, 0x64, 0x35, 0x34, 0x0a, 0x61, 0x75,\n+\t0x74, 0x68, 0x6f, 0x72, 0x20, 0x41, 0x20, 0x55,\n+\t0x20, 0x54, 0x68, 0x6f, 0x72, 0x20, 0x3c, 0x61,\n+\t0x75, 0x74, 0x68, 0x6f, 0x72, 0x40, 0x65, 0x78,\n+\t0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f,\n+\t0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33, 0x34, 0x35,\n+\t0x36, 0x37, 0x38, 0x39, 0x30, 0x20, 0x2b, 0x30,\n+\t0x30, 0x30, 0x30, 0x0a, 0x63, 0x6f, 0x6d, 0x6d,\n+\t0x69, 0x74, 0x74, 0x65, 0x72, 0x20, 0x43, 0x20,\n+\t0x4f, 0x20, 0x4d, 0x69, 0x74, 0x74, 0x65, 0x72,\n+\t0x20, 0x3c, 0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74,\n+\t0x74, 0x65, 0x72, 0x40, 0x65, 0x78, 0x61, 0x6d,\n+\t0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e,\n+\t0x20, 0x31, 0x32, 0x33, 0x34, 0x35, 0x36, 0x37,\n+\t0x38, 0x39, 0x30, 0x20, 0x2b, 0x30, 0x30, 0x30,\n+\t0x30, 0x0a, 0x0a, 0x45, 0x6d, 0x70, 0x74, 0x79,\n+\t0x20, 0x74, 0x72, 0x65, 0x65, 0x20, 0x63, 0x6f,\n+\t0x6d, 0x6d, 0x69, 0x74, 0x0a, 0x6d, 0x6d, 0x51,\n+\t0x9a, 0xc9, 0x11, 0x76, 0x61, 0xa3, 0x89, 0x49,\n+\t0xb7, 0xa1, 0x58, 0xc6, 0x1d, 0x8c, 0x33, 0x75,\n+\t0x8d, 0x7e, 0x4d, 0x8e, 0x58, 0x91, 0xf8, 0x5c,\n+\t0x57, 0xd9, 0x89, 0x9e, 0xb8, 0xd2, 0x9a, 0xd8,\n+\t0xc9\n+};\n+\n static const struct fast_pack fast_packs[] = {\n \t{\n \t\t.format_id = GIT_SHA1_FORMAT_ID,\n@@ -253,6 +337,13 @@ static const struct fast_pack fast_packs[] = {\n \t\t.suffix_len = sizeof(fast_pack_sha1_suffix),\n \t\t.commit_oid = \"aac43daf40d0377af31aa9c798a4ae8a31b55c1d\",\n \t},\n+\t{\n+\t\t.format_id = GIT_SHA256_FORMAT_ID,\n+\t\t.suffix = fast_pack_sha256_suffix,\n+\t\t.suffix_len = sizeof(fast_pack_sha256_suffix),\n+\t\t.commit_oid = \"63c46ca51267b1d45be69a044bb84b4bf0559f09\"\n+\t\t\t      \"d727f861d2ae94ddebdddbc9\",\n+\t},\n };\n \n /*\n-- \ngitgitgadget\n\n"},{"id":"542695","messageId":"5b44410b2f9b3ebf01d582445542b6aca9984c2e.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 10/11] t5608: mark >4GB tests as EXPENSIVE","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:27Z","receivedAt":"2026-05-04T17:08:43Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nEven with precomputed pack constants that reduced the helper's\nruntime from minutes to seconds, the >4GB clone tests still take\n200-850 seconds across CI jobs. The bottleneck is no longer the\npack generation but the clone operations themselves: transporting,\nunpacking, and indexing 4 GiB of data through unpack-objects and\nindex-pack is inherently expensive.\n\nAs Jeff King pointed out [1], t5608 alone takes 160 seconds on his\nlaptop while the rest of the entire test suite finishes in under 90\nseconds, and the test's disk footprint (4+ GiB source repo, then\ntwo clones) is problematic for developers who use RAM disks for\ntheir trash directories.\n\nGate the >4GB tests on the EXPENSIVE prereq (which requires\nGIT_TEST_LONG to be set) in addition to SIZE_T_IS_64BIT, keeping\nthem out of normal local test runs.\n\n[1] https://lore.kernel.org/git/20260501063805.GA2038915@coredump.intra.peff.net/\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t5608-clone-2gb.sh | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t5608-clone-2gb.sh b/t/t5608-clone-2gb.sh\nindex af93302dde..4f8a95ddda 100755\n--- a/t/t5608-clone-2gb.sh\n+++ b/t/t5608-clone-2gb.sh\n@@ -49,7 +49,7 @@ test_expect_success 'clone - with worktree, file:// protocol' '\n \n '\n \n-test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n+test_expect_success SIZE_T_IS_64BIT,EXPENSIVE 'set up repo with >4GB object' '\n \tlarge_blob_size=$((4*1024*1024*1024+1)) &&\n \tgit init --bare 4gb-repo &&\n \thead_oid=$(test-tool synthesize pack \\\n@@ -60,7 +60,7 @@ test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n \tgit -C 4gb-repo symbolic-ref HEAD refs/heads/main\n '\n \n-test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n+test_expect_success SIZE_T_IS_64BIT,EXPENSIVE 'clone >4GB object via unpack-objects' '\n \t# The synthesized pack has five objects, so a large unpack limit keeps\n \t# fetch-pack on the unpack-objects path.\n \tgit -c fetch.unpackLimit=100 clone --bare \\\n@@ -76,7 +76,7 @@ test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n \ttest \"$source_blob\" = \"$clone_blob\"\n '\n \n-test_expect_success SIZE_T_IS_64BIT 'clone with >4GB object via index-pack' '\n+test_expect_success SIZE_T_IS_64BIT,EXPENSIVE 'clone with >4GB object via index-pack' '\n \t# Force fetch-pack to hand the pack to index-pack instead.\n \tgit -c fetch.unpackLimit=1 clone --bare \\\n \t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-index &&\n-- \ngitgitgadget\n\n"},{"id":"542694","messageId":"1eaaa7fad7a1432dd97ffdd7c45e8162f61bc302.1777914508.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-04T17:08:28Z","receivedAt":"2026-05-04T17:08:44Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nDerrick Stolee suggested [1] that expensive tests should be run at a\nregular cadence rather than on every PR iteration. Gate GIT_TEST_LONG\non push builds to the integration branches (next, master, main, maint)\nso that the EXPENSIVE prereq is satisfied there but not during PR\nvalidation, where the extra minutes of wall-clock time do not justify\nthemselves.\n\n[1] https://lore.kernel.org/git/e1e8837f-7374-4079-ba87-ab95dd156e33@gmail.com/\n\nHelped-by: Derrick Stolee <derrickstolee@github.com>\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n ci/lib.sh | 9 +++++++++\n 1 file changed, 9 insertions(+)\n\ndiff --git a/ci/lib.sh b/ci/lib.sh\nindex 42a2b6a318..a671994bdf 100755\n--- a/ci/lib.sh\n+++ b/ci/lib.sh\n@@ -314,6 +314,15 @@ export DEFAULT_TEST_TARGET=prove\n export GIT_TEST_CLONE_2GB=true\n export SKIP_DASHED_BUILT_INS=YesPlease\n \n+# Enable expensive tests on push builds to integration branches, but\n+# not on PR builds where the extra time is not justified for every\n+# iteration.\n+case \"$GITHUB_EVENT_NAME,$CI_BRANCH\" in\n+push,*next*|push,*master*|push,*main*|push,*maint*)\n+\texport GIT_TEST_LONG=YesPlease\n+\t;;\n+esac\n+\n case \"$distro\" in\n ubuntu-*)\n \t# Python 2 is end of life, and Ubuntu 23.04 and newer don't actually\n-- \ngitgitgadget\n"},{"id":"542708","messageId":"a382fcdf-a9c9-4caa-8be4-163c7bcbd64b@gmail.com","threadId":"65563","inReplyTo":"29b9a74e915e6200ac2b4d98e446c1e73964cbd2.1777914508.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 08/11] test-tool synthesize: precompute pack for 4 GiB + 1","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-05-04T18:27:23Z","receivedAt":"2026-05-04T18:27:28Z","isPatch":true,"body":"On 5/4/2026 1:08 PM, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\n> Benchmarks generating a 4 GiB + 1 pack (3 runs each, SHA1DC on\n> x86_64):\n> \n>   generic path:   88s / 81s / 140s\n>   fast path:      14s / 13s / 15s\n> \n> On CI, where t5608 currently takes 200-850 seconds depending on the\n> job, the fast path cuts the pack-generation phase from minutes to\n> seconds, leaving only the clone operations themselves.\n\nAre these numbers accurate for the patch position in the series?\n\nThe previous change replaced SHA1DC with the unsafe version, which\ngained similar performance improvements. I'd be interested to see\nthe numbers for both enabled at the same time.\n\nThanks,\n-Stolee\n\n"},{"id":"542717","messageId":"42f96e54-7b94-4075-91b1-1c2447b93322@gmail.com","threadId":"65563","inReplyTo":"1eaaa7fad7a1432dd97ffdd7c45e8162f61bc302.1777914508.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-05-04T18:35:05Z","receivedAt":"2026-05-04T18:35:07Z","isPatch":true,"body":"On 5/4/2026 1:08 PM, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> Derrick Stolee suggested [1] that expensive tests should be run at a\n> regular cadence rather than on every PR iteration. Gate GIT_TEST_LONG\n> on push builds to the integration branches (next, master, main, maint)\n> so that the EXPENSIVE prereq is satisfied there but not during PR\n> validation, where the extra minutes of wall-clock time do not justify\n> themselves.\nI like that this will be run as part of regular updates to the\nimportant branches. The important bit after that is whether or\nnot a human pays attention to the signal of these builds.\n\nJunio: Do you pay attention to CI breaks when you push to\n'master'?\n\nOne way to help this procedure could be to have GitHub CI\nfailures trigger new issues, which could then be more easily\nviewed and noticed by the community watching the repo. This\nis of course out-of-scope for this patch series, but could be\nconsidered in the future.\n\nThanks,\n-Stolee\n\n"},{"id":"542760","messageId":"xmqq5x52nhg6.fsf@gitster.g","threadId":"65563","inReplyTo":"42f96e54-7b94-4075-91b1-1c2447b93322@gmail.com","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-05T12:56:09Z","receivedAt":"2026-05-05T12:56:12Z","isPatch":true,"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 5/4/2026 1:08 PM, Johannes Schindelin via GitGitGadget wrote:\n>> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>> \n>> Derrick Stolee suggested [1] that expensive tests should be run at a\n>> regular cadence rather than on every PR iteration. Gate GIT_TEST_LONG\n>> on push builds to the integration branches (next, master, main, maint)\n>> so that the EXPENSIVE prereq is satisfied there but not during PR\n>> validation, where the extra minutes of wall-clock time do not justify\n>> themselves.\n> I like that this will be run as part of regular updates to the\n> important branches. The important bit after that is whether or\n> not a human pays attention to the signal of these builds.\n>\n> Junio: Do you pay attention to CI breaks when you push to\n> 'master'?\n\nWell, it is way too late to notice breakage when the faulty update\nhits 'master'.  CI failures should be noticed before breakage hits\n'next'.\n\nI often notice and complain when I see failures on 'seen', and\nsometimes I help original submitter by bisecting, but I do not\nnecessarily have enough time and bandwidth to help everybody.\n\nQuite honestly, the best place to give widest test coverage is much\ncloser to the source of the problems than in my tree and mixed with\nother topics, i.e., at individual contributor's CI.  That way, I\npresume that GitGitGadget can also help submitters avoid sending a\nfaulty series, reducing the load on the list and the maintainer.\n\nIdeally the CI tests by the integrator should only be catching any\nmismerges and unexpected inter-topic interactions, as they cannot be\ncaught by contributor's standalone tests, so I do not mind widening\ncoverage of CI tests when I push the integration results out.  But\nso far, the majority of what I have seen and reported back to the\nlist have been something that the authors should be equipped to spot\nin their topic without getting mixed with other topics into any\nintegration branches.\n\n> One way to help this procedure could be to have GitHub CI\n> failures trigger new issues, which could then be more easily\n> viewed and noticed by the community watching the repo. This\n> is of course out-of-scope for this patch series, but could be\n> considered in the future.\n\nI think a better way to help would be to arrange the workflow so\nthat we do not even have to trigger an issue, and stop before the\npatches leave the original authors' hand.  They can of course ask\nfor help saying \"here is my topic in my fork of the repository and\nfailing in this way for macOS that I do not have access to.  Could\nanybody help me figuring out what macOS peculiarity my changes are\ntickling?\", or something like that.\n\nIt would be best to find problems early, and make it easier for\nindividual contributors to help each other by having a concrete CI\nfailure reports in their forks that they can point at when they ask\nfor help.  And CI run when I push 'seen' or 'master' out would not\nhelp as much as CI run when they publish their forked branches would.\n\nBy the way, please expect slow responses as I am (officially) still\nmostly offline for the rest of the week.\n\nThanks.\n\n"},{"id":"542768","messageId":"20260505191100.GA12275@tb-raspi4","threadId":"65563","inReplyTo":"dc660106ea8511e6adc44d2b70e9a4ae8b18090e.1777914508.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-05-05T19:11:00Z","receivedAt":"2026-05-05T19:16:27Z","isPatch":true,"body":"On Mon, May 04, 2026 at 05:08:18PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> When unpacking objects from a packfile, the object size is decoded\n> from a variable-length encoding. On platforms where unsigned long is\n> 32-bit (such as Windows, even in 64-bit builds), the shift operation\n> overflows when decoding sizes larger than 4GB. The result is a\n> truncated size value, causing the unpacked object to be corrupted or\n> rejected.\n> \n> Fix this by changing the size variable to size_t, which is 64-bit on\n> 64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\n> space.\n> \n> This was originally authored by LordKiRon <https://github.com/LordKiRon>,\n> who preferred not to reveal their real name and therefore agreed that I\n> take over authorship.\n> \n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  builtin/index-pack.c     | 9 +++++----\n>  builtin/unpack-objects.c | 5 +++--\n>  2 files changed, 8 insertions(+), 6 deletions(-)\n> \n> diff --git a/builtin/index-pack.c b/builtin/index-pack.c\n> index ca7784dc2c..cc660582e9 100644\n> --- a/builtin/index-pack.c\n> +++ b/builtin/index-pack.c\n> @@ -37,7 +37,7 @@ static const char index_pack_usage[] =\n>  \n>  struct object_entry {\n>  \tstruct pack_idx_entry idx;\n> -\tunsigned long size;\n> +\tsize_t size;\n>  \tunsigned char hdr_size;\n>  \tsigned char type;\n>  \tsigned char real_type;\n> @@ -469,7 +469,7 @@ static int is_delta_type(enum object_type type)\n>  \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n>  }\n>  \n> -static void *unpack_entry_data(off_t offset, unsigned long size,\n> +static void *unpack_entry_data(off_t offset, size_t size,\n>  \t\t\t       enum object_type type, struct object_id *oid)\n>  {\n>  \tstatic char fixed_buf[8192];\n> @@ -524,7 +524,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n>  \t\t\t      struct object_id *oid)\n>  {\n>  \tunsigned char *p;\n> -\tunsigned long size, c;\n> +\tsize_t size;\n> +\tunsigned long c;\nDoes this look a little bit strange ?\np points to an unsigned char (better would be *uint8_t)\nthen it is dereferenced into an \"unsigned long\".\nThen it is masked with 0x7f\nIn short: should \"c\" be declared as uint8_t ?\n\n>  \toff_t base_offset;\n>  \tunsigned shift;\n>  \tvoid *data;\n> @@ -542,7 +543,7 @@ static void *unpack_raw_entry(struct object_entry *obj,\n>  \t\tp = fill(1);\n>  \t\tc = *p;\n>  \t\tuse(1);\n> -\t\tsize += (c & 0x7f) << shift;\n> +\t\tsize += ((size_t)c & 0x7f) << shift;\n>  \t\tshift += 7;\n>  \t}\n>  \tobj->size = size;\n> diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> index e01cf6e360..59a36c2481 100644\n> --- a/builtin/unpack-objects.c\n> +++ b/builtin/unpack-objects.c\n> @@ -533,7 +533,8 @@ static void unpack_one(unsigned nr)\n>  {\n>  \tunsigned shift;\n>  \tunsigned char *pack;\n> -\tunsigned long size, c;\n> +\tsize_t size;\n> +\tunsigned long c;\n>  \tenum object_type type;\n>  \n>  \tobj_list[nr].offset = consumed_bytes;\n> @@ -548,7 +549,7 @@ static void unpack_one(unsigned nr)\n>  \t\tpack = fill(1);\n>  \t\tc = *pack;\n>  \t\tuse(1);\n> -\t\tsize += (c & 0x7f) << shift;\n> +\t\tsize += ((size_t)c & 0x7f) << shift;\n>  \t\tshift += 7;\n>  \t}\n>  \n> -- \n> gitgitgadget\n> \n> \n"},{"id":"542777","messageId":"20260505192722.GB12275@tb-raspi4","threadId":"65563","inReplyTo":"3a539061c5f62c65d46bd0eb774bb1b1239463ff.1777914508.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 03/11] odb, packfile: use size_t for streaming object sizes","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-05-05T19:27:22Z","receivedAt":"2026-05-05T19:27:25Z","isPatch":true,"body":"On Mon, May 04, 2026 at 05:08:20PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> The odb_read_stream structure uses unsigned long for the size field,\n> which is 32-bit on Windows even in 64-bit builds. When streaming\n> objects larger than 4GB, the size would be truncated to zero or an\n> incorrect value, resulting in empty files being written to disk.\n> \n> Change the size field in odb_read_stream to size_t and introduce\n> unpack_object_header_sz() to return sizes via size_t pointer. Since\n> object_info.sizep remains unsigned long for API compatibility, use\n> temporary variables where the types differ, with comments noting the\n> truncation limitation for code paths that still use unsigned long.\n> \n> This was originally authored by LordKiRon <https://github.com/LordKiRon>,\n> who preferred not to reveal their real name and therefore agreed that I\n> take over authorship.\n> \n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n>  builtin/pack-objects.c       | 23 ++++++++++++++++-------\n>  object-file.c                | 10 +++++++++-\n>  odb/streaming.c              | 13 ++++++++++++-\n>  odb/streaming.h              |  2 +-\n>  oss-fuzz/fuzz-pack-headers.c |  2 +-\n>  pack-bitmap.c                |  2 +-\n>  pack-check.c                 |  6 ++++--\n>  packfile.c                   | 24 +++++++++++++++---------\n>  packfile.h                   |  4 ++--\n>  9 files changed, 61 insertions(+), 25 deletions(-)\n> \n\n> diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> index dd2480a73d..aa4b1cb9b8 100644\n> --- a/builtin/pack-objects.c\n> +++ b/builtin/pack-objects.c\n\nI haven't been able to follow all changes, so this may be false alarm.\nDo we need a cast_size_t_to_ulong() somewhere ?\n\n> @@ -629,14 +629,21 @@ static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n>  \tstruct packed_git *p = IN_PACK(entry);\n>  \tstruct pack_window *w_curs = NULL;\n>  \tuint32_t pos;\n> -\toff_t offset;\n> +\toff_t offset, cur;\n>  \tenum object_type type = oe_type(entry);\n> +\tenum object_type in_pack_type;\n>  \toff_t datalen;\n>  \tunsigned char header[MAX_PACK_OBJECT_HEADER],\n>  \t\t      dheader[MAX_PACK_OBJECT_HEADER];\n>  \tunsigned hdrlen;\n>  \tconst unsigned hashsz = the_hash_algo->rawsz;\n> -\tunsigned long entry_size = SIZE(entry);\n> +\tsize_t entry_size;\n> +\n> +\tcur = entry->in_pack_offset;\n> +\tin_pack_type = unpack_object_header(p, &w_curs, &cur, &entry_size);\n> +\tif (in_pack_type < 0)\n> +\t\tdie(_(\"write_reuse_object: unable to parse object header of %s\"),\n> +\t\t    oid_to_hex(&entry->idx.oid));\n>  \n>  \tif (DELTA(entry))\n>  \t\ttype = (allow_ofs_delta && DELTA(entry)->idx.offset) ?\n> @@ -1087,7 +1094,7 @@ static void write_reused_pack_one(struct packed_git *reuse_packfile,\n>  {\n>  \toff_t offset, next, cur;\n>  \tenum object_type type;\n> -\tunsigned long size;\n> +\tsize_t size;\n>  \n>  \toffset = pack_pos_to_offset(reuse_packfile, pos);\n>  \tnext = pack_pos_to_offset(reuse_packfile, pos + 1);\n> @@ -2243,7 +2250,7 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n>  \t\toff_t ofs;\n>  \t\tunsigned char *buf, c;\n>  \t\tenum object_type type;\n> -\t\tunsigned long in_pack_size;\n> +\t\tsize_t in_pack_size;\n>  \n>  \t\tbuf = use_pack(p, &w_curs, entry->in_pack_offset, &avail);\n>  \n> @@ -2734,16 +2741,18 @@ unsigned long oe_get_size_slow(struct packing_data *pack,\n>  \tstruct pack_window *w_curs;\n>  \tunsigned char *buf;\n>  \tenum object_type type;\n> -\tunsigned long used, avail, size;\n> +\tunsigned long used, avail;\n> +\tsize_t size;\n>  \n>  \tif (e->type_ != OBJ_OFS_DELTA && e->type_ != OBJ_REF_DELTA) {\n> +\t\tunsigned long sz;\n>  \t\tpacking_data_lock(&to_pack);\n>  \t\tif (odb_read_object_info(the_repository->objects,\n> -\t\t\t\t\t &e->idx.oid, &size) < 0)\n> +\t\t\t\t\t &e->idx.oid, &sz) < 0)\n>  \t\t\tdie(_(\"unable to get size of %s\"),\n>  \t\t\t    oid_to_hex(&e->idx.oid));\n>  \t\tpacking_data_unlock(&to_pack);\n> -\t\treturn size;\n> +\t\treturn sz;\n>  \t}\n>  \n>  \tp = oe_in_pack(pack, e);\n> diff --git a/object-file.c b/object-file.c\n> index 086b2b65ff..0be2981c7a 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -2326,6 +2326,7 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n>  \tstruct object_info oi = OBJECT_INFO_INIT;\n>  \tstruct odb_loose_read_stream *st;\n>  \tunsigned long mapsize;\n> +\tunsigned long size_ul;\n>  \tvoid *mapped;\n>  \n>  \tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n> @@ -2349,11 +2350,18 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n>  \t\tgoto error;\n>  \t}\n>  \n> -\toi.sizep = &st->base.size;\n> +\t/*\n> +\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n> +\t * st->base.size is size_t (64-bit). Use temporary variable.\n> +\t * Note: loose objects >4GB would still truncate here, but such\n> +\t * large loose objects are uncommon (they'd normally be packed).\n> +\t */\n> +\toi.sizep = &size_ul;\n>  \toi.typep = &st->base.type;\n>  \n>  \tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n>  \t\tgoto error;\n> +\tst->base.size = size_ul;\n>  \n>  \tst->mapped = mapped;\n>  \tst->mapsize = mapsize;\n> diff --git a/odb/streaming.c b/odb/streaming.c\n> index 5927a12954..af2adf5ce7 100644\n> --- a/odb/streaming.c\n> +++ b/odb/streaming.c\n> @@ -157,15 +157,26 @@ static int open_istream_incore(struct odb_read_stream **out,\n>  \t\t.base.read = read_istream_incore,\n>  \t};\n>  \tstruct odb_incore_read_stream *st;\n> +\tunsigned long size_ul;\n>  \tint ret;\n>  \n>  \toi.typep = &stream.base.type;\n> -\toi.sizep = &stream.base.size;\n> +\t/*\n> +\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n> +\t * stream.base.size is size_t (64-bit). We use a temporary variable\n> +\t * because the types are incompatible. Note: this path still truncates\n> +\t * for >4GB objects, but large objects should use pack streaming\n> +\t * (packfile_store_read_object_stream) which handles size_t properly.\n> +\t * This incore fallback is only used for small objects or when pack\n> +\t * streaming is unavailable.\n> +\t */\n> +\toi.sizep = &size_ul;\n>  \toi.contentp = (void **)&stream.buf;\n>  \tret = odb_read_object_info_extended(odb, oid, &oi,\n>  \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n>  \tif (ret)\n>  \t\treturn ret;\n> +\tstream.base.size = size_ul;\n>  \n>  \tCALLOC_ARRAY(st, 1);\n>  \t*st = stream;\n> diff --git a/odb/streaming.h b/odb/streaming.h\n> index c7861f7e13..517e2ea2d3 100644\n> --- a/odb/streaming.h\n> +++ b/odb/streaming.h\n> @@ -21,7 +21,7 @@ struct odb_read_stream {\n>  \todb_read_stream_close_fn close;\n>  \todb_read_stream_read_fn read;\n>  \tenum object_type type;\n> -\tunsigned long size; /* inflated size of full object */\n> +\tsize_t size; /* inflated size of full object */\n>  };\n>  \n>  /*\n> diff --git a/oss-fuzz/fuzz-pack-headers.c b/oss-fuzz/fuzz-pack-headers.c\n> index 150c0f5fa2..ef61ab577c 100644\n> --- a/oss-fuzz/fuzz-pack-headers.c\n> +++ b/oss-fuzz/fuzz-pack-headers.c\n> @@ -6,7 +6,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size);\n>  int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n>  {\n>  \tenum object_type type;\n> -\tunsigned long len;\n> +\tsize_t len;\n>  \n>  \tunpack_object_header_buffer((const unsigned char *)data,\n>  \t\t\t\t    (unsigned long)size, &type, &len);\n> diff --git a/pack-bitmap.c b/pack-bitmap.c\n> index f6ec18d83a..f9af8a96bd 100644\n> --- a/pack-bitmap.c\n> +++ b/pack-bitmap.c\n> @@ -2270,7 +2270,7 @@ static int try_partial_reuse(struct bitmap_index *bitmap_git,\n>  {\n>  \toff_t delta_obj_offset;\n>  \tenum object_type type;\n> -\tunsigned long size;\n> +\tsize_t size;\n>  \n>  \tif (pack_pos >= pack->p->num_objects)\n>  \t\treturn -1; /* not actually in the pack */\n> diff --git a/pack-check.c b/pack-check.c\n> index 79992bb509..2792f34d25 100644\n> --- a/pack-check.c\n> +++ b/pack-check.c\n> @@ -110,7 +110,7 @@ static int verify_packfile(struct repository *r,\n>  \t\tvoid *data;\n>  \t\tstruct object_id oid;\n>  \t\tenum object_type type;\n> -\t\tunsigned long size;\n> +\t\tsize_t size;\n>  \t\toff_t curpos;\n>  \t\tint data_valid;\n>  \n> @@ -143,7 +143,9 @@ static int verify_packfile(struct repository *r,\n>  \t\t\tdata = NULL;\n>  \t\t\tdata_valid = 0;\n>  \t\t} else {\n> -\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &size);\n> +\t\t\tunsigned long sz;\n> +\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &sz);\n> +\t\t\tsize = sz;\n>  \t\t\tdata_valid = 1;\n>  \t\t}\n>  \n> diff --git a/packfile.c b/packfile.c\n> index b012d648ad..fdae91dd11 100644\n> --- a/packfile.c\n> +++ b/packfile.c\n> @@ -1133,7 +1133,7 @@ out:\n>  }\n>  \n>  unsigned long unpack_object_header_buffer(const unsigned char *buf,\n> -\t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n> +\t\tunsigned long len, enum object_type *type, size_t *sizep)\n>  {\n>  \tunsigned shift;\n>  \tsize_t size, c;\n> @@ -1144,7 +1144,11 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n>  \tsize = c & 15;\n>  \tshift = 4;\n>  \twhile (c & 0x80) {\n> -\t\tif (len <= used || (bitsizeof(long) - 7) < shift) {\n> +\t\t/*\n> +\t\t * Each continuation byte adds 7 bits. Ensure shift won't\n> +\t\t * overflow size_t (use size_t not long for 64-bit on Windows).\n> +\t\t */\n> +\t\tif (len <= used || (bitsizeof(size_t) - 7) < shift) {\n>  \t\t\terror(\"bad object header\");\n>  \t\t\tsize = used = 0;\n>  \t\t\tbreak;\n> @@ -1153,7 +1157,7 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n>  \t\tsize = st_add(size, st_left_shift(c & 0x7f, shift));\n>  \t\tshift += 7;\n>  \t}\n> -\t*sizep = cast_size_t_to_ulong(size);\n> +\t*sizep = size;\n>  \treturn used;\n>  }\n>  \n> @@ -1215,7 +1219,7 @@ unsigned long get_size_from_delta(struct packed_git *p,\n>  int unpack_object_header(struct packed_git *p,\n>  \t\t\t struct pack_window **w_curs,\n>  \t\t\t off_t *curpos,\n> -\t\t\t unsigned long *sizep)\n> +\t\t\t size_t *sizep)\n>  {\n>  \tunsigned char *base;\n>  \tunsigned long left;\n> @@ -1367,7 +1371,7 @@ static enum object_type packed_to_object_type(struct repository *r,\n>  \n>  \twhile (type == OBJ_OFS_DELTA || type == OBJ_REF_DELTA) {\n>  \t\toff_t base_offset;\n> -\t\tunsigned long size;\n> +\t\tsize_t size;\n>  \t\t/* Push the object we're going to leave behind */\n>  \t\tif (poi_stack_nr >= poi_stack_alloc && poi_stack == small_poi_stack) {\n>  \t\t\tpoi_stack_alloc = alloc_nr(poi_stack_nr);\n> @@ -1586,7 +1590,7 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n>  \t\t\t\t\t     uint32_t *maybe_index_pos, struct object_info *oi)\n>  {\n>  \tstruct pack_window *w_curs = NULL;\n> -\tunsigned long size;\n> +\tsize_t size;\n>  \toff_t curpos = obj_offset;\n>  \tenum object_type type = OBJ_NONE;\n>  \tuint32_t pack_pos;\n> @@ -1778,7 +1782,7 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n>  \tstruct pack_window *w_curs = NULL;\n>  \toff_t curpos = obj_offset;\n>  \tvoid *data = NULL;\n> -\tunsigned long size;\n> +\tsize_t size;\n>  \tenum object_type type;\n>  \tstruct unpack_entry_stack_ent small_delta_stack[UNPACK_ENTRY_STACK_PREALLOC];\n>  \tstruct unpack_entry_stack_ent *delta_stack = small_delta_stack;\n> @@ -1943,8 +1947,10 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n>  \t\t\t      (uintmax_t)curpos, p->pack_name);\n>  \t\t\tdata = NULL;\n>  \t\t} else {\n> +\t\t\tunsigned long sz;\n>  \t\t\tdata = patch_delta(base, base_size, delta_data,\n> -\t\t\t\t\t   delta_size, &size);\n> +\t\t\t\t\t   delta_size, &sz);\n> +\t\t\tsize = sz;\n>  \n>  \t\t\t/*\n>  \t\t\t * We could not apply the delta; warn the user, but\n> @@ -2929,7 +2935,7 @@ int packfile_read_object_stream(struct odb_read_stream **out,\n>  \tstruct odb_packed_read_stream *stream;\n>  \tstruct pack_window *window = NULL;\n>  \tenum object_type in_pack_type;\n> -\tunsigned long size;\n> +\tsize_t size;\n>  \n>  \tin_pack_type = unpack_object_header(pack, &window, &offset, &size);\n>  \tunuse_pack(&window);\n> diff --git a/packfile.h b/packfile.h\n> index 9b647da7dd..49d6bdecf6 100644\n> --- a/packfile.h\n> +++ b/packfile.h\n> @@ -456,9 +456,9 @@ off_t find_pack_entry_one(const struct object_id *oid, struct packed_git *);\n>  \n>  int is_pack_valid(struct packed_git *);\n>  void *unpack_entry(struct repository *r, struct packed_git *, off_t, enum object_type *, unsigned long *);\n> -unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n> +unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, size_t *sizep);\n>  unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n> -int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\n> +int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, size_t *);\n>  off_t get_delta_base(struct packed_git *p, struct pack_window **w_curs,\n>  \t\t     off_t *curpos, enum object_type type,\n>  \t\t     off_t delta_obj_offset);\n> -- \n> gitgitgadget\n> \n> \n"},{"id":"542787","messageId":"53431a2a-a0e8-1dd1-9ebd-cbaebc769a39@gmx.de","threadId":"65563","inReplyTo":"a382fcdf-a9c9-4caa-8be4-163c7bcbd64b@gmail.com","subject":"Re: [PATCH v2 08/11] test-tool synthesize: precompute pack for 4 GiB + 1","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-05T20:54:59Z","receivedAt":"2026-05-05T21:00:13Z","isPatch":true,"body":"Hi Stolee,\n\nOn Tue, 5 May 2026, Derrick Stolee wrote:\n\n> On 5/4/2026 1:08 PM, Johannes Schindelin via GitGitGadget wrote:\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> \n> > Benchmarks generating a 4 GiB + 1 pack (3 runs each, SHA1DC on\n> > x86_64):\n> > \n> >   generic path:   88s / 81s / 140s\n> >   fast path:      14s / 13s / 15s\n> > \n> > On CI, where t5608 currently takes 200-850 seconds depending on the\n> > job, the fast path cuts the pack-generation phase from minutes to\n> > seconds, leaving only the clone operations themselves.\n> \n> Are these numbers accurate for the patch position in the series?\n\nUnfortunately, yes.\n\n> The previous change replaced SHA1DC with the unsafe version, which\n> gained similar performance improvements. I'd be interested to see\n> the numbers for both enabled at the same time.\n\nThe problem is that in many (most?) cases, the \"unsafe\" version is the\n_same_ as the safe version, i.e. SHA1DC. Only the `linux-TEST-vars` job on\nCI (and most notably, _not_ in the `win-*` jobs) has a fast \"unsafe\"\nversion by default. So if I build the revision as per the previous patch\non Windows, I get no performance benefit whatsoever. That's what my lament\nabout not being able to link OpenSSL was all about.\n\nCiao,\nJohannes\n"},{"id":"542796","messageId":"CAPc5daUzr+mn6ojzsqpW6mCXzc2yVqpevVk8njefx4j09G_OgA@mail.gmail.com","threadId":"65563","inReplyTo":"xmqq5x52nhg6.fsf@gitster.g","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-05T23:07:27Z","receivedAt":"2026-05-05T23:07:47Z","isPatch":true,"body":"(in GMail web interface, excuse typos)\n\nhttps://github.com/git/git/actions/runs/25366120610/job/74377320625\n\nWe seem to be hitting the same _Generic error in various (but not all) jobs\n\n  /usr/include/x86_64-linux-gnu/sys/cdefs.h:838:3: note: expanded from\nmacro '__glibc_const_generic'\n    838 |   _Generic (0 ? (PTR) : (void *) 1,                     \\\n        |   ^\n  Error: list-objects-filter-options.c:222:10: '_Generic' is a C11\nextension [-Werror,-Wc11-extensions]\n\nI thought we updated the codebase to avoid stripping away constness\nwith strchr() and friends, but the error seems to be more like one\nhand in the system passing -Wc11-extensions to stick to older version\nof C and the other hand in the system that uses _Generic to implement\nthe const/non-const variants of strchr() in the system header not\nknowing that the other tells C11 const-preserving strchr() should not\nbe used?\n\n\n2026年5月5日(火) 21:56 Junio C Hamano <gitster@pobox.com>:\n>\n> Derrick Stolee <stolee@gmail.com> writes:\n>\n> > On 5/4/2026 1:08 PM, Johannes Schindelin via GitGitGadget wrote:\n> >> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> >>\n> >> Derrick Stolee suggested [1] that expensive tests should be run at a\n> >> regular cadence rather than on every PR iteration. Gate GIT_TEST_LONG\n> >> on push builds to the integration branches (next, master, main, maint)\n> >> so that the EXPENSIVE prereq is satisfied there but not during PR\n> >> validation, where the extra minutes of wall-clock time do not justify\n> >> themselves.\n> > I like that this will be run as part of regular updates to the\n> > important branches. The important bit after that is whether or\n> > not a human pays attention to the signal of these builds.\n> >\n> > Junio: Do you pay attention to CI breaks when you push to\n> > 'master'?\n>\n> Well, it is way too late to notice breakage when the faulty update\n> hits 'master'.  CI failures should be noticed before breakage hits\n> 'next'.\n>\n> I often notice and complain when I see failures on 'seen', and\n> sometimes I help original submitter by bisecting, but I do not\n> necessarily have enough time and bandwidth to help everybody.\n>\n> Quite honestly, the best place to give widest test coverage is much\n> closer to the source of the problems than in my tree and mixed with\n> other topics, i.e., at individual contributor's CI.  That way, I\n> presume that GitGitGadget can also help submitters avoid sending a\n> faulty series, reducing the load on the list and the maintainer.\n>\n> Ideally the CI tests by the integrator should only be catching any\n> mismerges and unexpected inter-topic interactions, as they cannot be\n> caught by contributor's standalone tests, so I do not mind widening\n> coverage of CI tests when I push the integration results out.  But\n> so far, the majority of what I have seen and reported back to the\n> list have been something that the authors should be equipped to spot\n> in their topic without getting mixed with other topics into any\n> integration branches.\n>\n> > One way to help this procedure could be to have GitHub CI\n> > failures trigger new issues, which could then be more easily\n> > viewed and noticed by the community watching the repo. This\n> > is of course out-of-scope for this patch series, but could be\n> > considered in the future.\n>\n> I think a better way to help would be to arrange the workflow so\n> that we do not even have to trigger an issue, and stop before the\n> patches leave the original authors' hand.  They can of course ask\n> for help saying \"here is my topic in my fork of the repository and\n> failing in this way for macOS that I do not have access to.  Could\n> anybody help me figuring out what macOS peculiarity my changes are\n> tickling?\", or something like that.\n>\n> It would be best to find problems early, and make it easier for\n> individual contributors to help each other by having a concrete CI\n> failure reports in their forks that they can point at when they ask\n> for help.  And CI run when I push 'seen' or 'master' out would not\n> help as much as CI run when they publish their forked branches would.\n>\n> By the way, please expect slow responses as I am (officially) still\n> mostly offline for the rest of the week.\n>\n> Thanks.\n>\n"},{"id":"542802","messageId":"e00dbf04-5866-008f-12e9-efdaacc3f2e0@gmx.de","threadId":"65563","inReplyTo":"CAPc5daUzr+mn6ojzsqpW6mCXzc2yVqpevVk8njefx4j09G_OgA@mail.gmail.com","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-06T08:33:56Z","receivedAt":"2026-05-06T08:34:10Z","isPatch":true,"body":"Hi Junio,\n\nOn Wed, 6 May 2026, Junio C Hamano wrote:\n\n> https://github.com/git/git/actions/runs/25366120610/job/74377320625\n> \n> We seem to be hitting the same _Generic error in various (but not all) jobs\n> \n>   /usr/include/x86_64-linux-gnu/sys/cdefs.h:838:3: note: expanded from\n> macro '__glibc_const_generic'\n>     838 |   _Generic (0 ? (PTR) : (void *) 1,                     \\\n>         |   ^\n>   Error: list-objects-filter-options.c:222:10: '_Generic' is a C11\n> extension [-Werror,-Wc11-extensions]\n> \n> I thought we updated the codebase to avoid stripping away constness\n> with strchr() and friends, but the error seems to be more like one\n> hand in the system passing -Wc11-extensions to stick to older version\n> of C and the other hand in the system that uses _Generic to implement\n> the const/non-const variants of strchr() in the system header not\n> knowing that the other tells C11 const-preserving strchr() should not\n> be used?\n\nThis was diagnosed (with a proposed fix) by Patrick over in\nhttps://lore.kernel.org/git/20260505-b4-pks-ci-tolerate-glibc-generic-v1-1-5786386fe512@pks.im/.\n\ntl;dr It's not about `const`-ness at all, but about glibc using a C11\nconstruct which clang's strict c99 checker now refuses, thanks to the\nupgrade to Ubuntu 26.04 in the `ubuntu:rolling` runners.\n\nCiao,\nJohannes\n"},{"id":"542835","messageId":"87se83efx1.fsf@gitster.g","threadId":"65563","inReplyTo":"e00dbf04-5866-008f-12e9-efdaacc3f2e0@gmx.de","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-07T09:18:34Z","receivedAt":"2026-05-07T09:18:41Z","isPatch":true,"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n>> I thought we updated the codebase to avoid stripping away constness\n>> with strchr() and friends, but the error seems to be more like one\n>> hand in the system passing -Wc11-extensions to stick to older version\n>> of C and the other hand in the system that uses _Generic to implement\n>> the const/non-const variants of strchr() in the system header not\n>> knowing that the other tells C11 const-preserving strchr() should not\n>> be used?\n>\n> This was diagnosed (with a proposed fix) by Patrick over in\n> https://lore.kernel.org/git/20260505-b4-pks-ci-tolerate-glibc-generic-v1-1-5786386fe512@pks.im/.\n\nIndeed.\n\n> tl;dr It's not about `const`-ness at all, but about glibc using a C11\n> construct which clang's strict c99 checker now refuses, thanks to the\n> upgrade to Ubuntu 26.04 in the `ubuntu:rolling` runners.\n\nYes, that is exactly what I meant by one hand knowing that it was\ntold not to use c11 extensions while the other hand ignoring and\nalways using c11 extensions in the header.  I recall that in the\npast gnu library headers were a bit more careful to make the life\nmore pleasant when we use (or decline to use) various features by\nusing conditional compilation, but apparently not this case.\n"},{"id":"542839","messageId":"afxoQh8SxCqBCaFP@pks.im","threadId":"65563","inReplyTo":"87se83efx1.fsf@gitster.g","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-05-07T10:24:02Z","receivedAt":"2026-05-07T10:24:10Z","isPatch":true,"body":"On Thu, May 07, 2026 at 06:18:34PM +0900, Junio C Hamano wrote:\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> >> I thought we updated the codebase to avoid stripping away constness\n> >> with strchr() and friends, but the error seems to be more like one\n> >> hand in the system passing -Wc11-extensions to stick to older version\n> >> of C and the other hand in the system that uses _Generic to implement\n> >> the const/non-const variants of strchr() in the system header not\n> >> knowing that the other tells C11 const-preserving strchr() should not\n> >> be used?\n> >\n> > This was diagnosed (with a proposed fix) by Patrick over in\n> > https://lore.kernel.org/git/20260505-b4-pks-ci-tolerate-glibc-generic-v1-1-5786386fe512@pks.im/.\n> \n> Indeed.\n> \n> > tl;dr It's not about `const`-ness at all, but about glibc using a C11\n> > construct which clang's strict c99 checker now refuses, thanks to the\n> > upgrade to Ubuntu 26.04 in the `ubuntu:rolling` runners.\n> \n> Yes, that is exactly what I meant by one hand knowing that it was\n> told not to use c11 extensions while the other hand ignoring and\n> always using c11 extensions in the header.  I recall that in the\n> past gnu library headers were a bit more careful to make the life\n> more pleasant when we use (or decline to use) various features by\n> using conditional compilation, but apparently not this case.\n\nYeah, it's a bit unfortunate indeed. I'd claim that this is a plain bug\nthough -- as mentioned in the commit message, I think what glibc should\nhave used is `_has_feature()` instead of `_has_extension()`,  and if so,\nI think the issue wouldn't exist.\n\nBut oh, well.\n\nPatrick\n"},{"id":"542883","messageId":"xmqqqznmfwd1.fsf@gitster.g","threadId":"65563","inReplyTo":"xmqq5x52nhg6.fsf@gitster.g","subject":"Re: [PATCH v2 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-08T02:50:18Z","receivedAt":"2026-05-08T02:50:21Z","isPatch":true,"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Derrick Stolee <stolee@gmail.com> writes:\n>\n>> On 5/4/2026 1:08 PM, Johannes Schindelin via GitGitGadget wrote:\n>>> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>>> \n>>> Derrick Stolee suggested [1] that expensive tests should be run at a\n>>> regular cadence rather than on every PR iteration. Gate GIT_TEST_LONG\n>>> on push builds to the integration branches (next, master, main, maint)\n>>> so that the EXPENSIVE prereq is satisfied there but not during PR\n>>> validation, where the extra minutes of wall-clock time do not justify\n>>> themselves.\n>> I like that this will be run as part of regular updates to the\n>> important branches. The important bit after that is whether or\n>> not a human pays attention to the signal of these builds.\n>>\n>> Junio: Do you pay attention to CI breaks when you push to\n>> 'master'?\n>\n> Well, it is way too late to notice breakage when the faulty update\n> hits 'master'.  CI failures should be noticed before breakage hits\n> 'next'.\n>\n> I often notice and complain when I see failures on 'seen', and\n> sometimes I help original submitter by bisecting, but I do not\n> necessarily have enough time and bandwidth to help everybody.\n\nTo more directly answer your question, yes I do pay attention, but\nnot only when I push to 'master', but when I push to any of the\nintegration branches.  I pay most attention to breakage in 'seen',\nso that I can notify authors of new topics early.  Sometimes you may\nsee many pushes only to 'seen' at github.com/git/git while 'seen' on\nthe other hosting sites are not updated with these commits, and if\nyou notice them, you caught me bisecting the breakage on it.  This\nis so that I can eject offending topics and notify the author.\n\nBut this does not scale, and I shouldn't have to do it myself.\n\nMaking sure the topics come in a shape that they pass the tests\nbefore they hit my tree is one way to reduce the need for the\nmaintainer bottleneck.\n\n> It would be best to find problems early, and make it easier for\n> individual contributors to help each other by having a concrete CI\n> failure reports in their forks that they can point at when they ask\n> for help.  And CI run when I push 'seen' or 'master' out would not\n> help as much as CI run when they publish their forked branches would.\n>\n> By the way, please expect slow responses as I am (officially) still\n> mostly offline for the rest of the week.\n>\n> Thanks.\n"},{"id":"542887","messageId":"fa39d84b-ddbc-3943-5cca-078fb18db80d@gmx.de","threadId":"65563","inReplyTo":"20260505191100.GA12275@tb-raspi4","subject":"Re: [PATCH v2 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-08T07:36:53Z","receivedAt":"2026-05-08T07:36:56Z","isPatch":true,"body":"Hi Torsten,\n\nOn Tue, 5 May 2026, Torsten Bögershausen wrote:\n\n> On Mon, May 04, 2026 at 05:08:18PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > \n> > [...]\n> > @@ -524,7 +524,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n> >  \t\t\t      struct object_id *oid)\n> >  {\n> >  \tunsigned char *p;\n> > -\tunsigned long size, c;\n> > +\tsize_t size;\n> > +\tunsigned long c;\n>\n> Does this look a little bit strange ?\n\nGood point.\n\n> p points to an unsigned char (better would be *uint8_t)\n> then it is dereferenced into an \"unsigned long\".\n> Then it is masked with 0x7f\n> In short: should \"c\" be declared as uint8_t ?\n\nAlmost. It should be a `size_t`, so that we don't have to cast it when\nshifting it. I'll include a fix in the next iteration.\n\nThank you!\nJohannes\n\n> \n> >  \toff_t base_offset;\n> >  \tunsigned shift;\n> >  \tvoid *data;\n> > @@ -542,7 +543,7 @@ static void *unpack_raw_entry(struct object_entry *obj,\n> >  \t\tp = fill(1);\n> >  \t\tc = *p;\n> >  \t\tuse(1);\n> > -\t\tsize += (c & 0x7f) << shift;\n> > +\t\tsize += ((size_t)c & 0x7f) << shift;\n> >  \t\tshift += 7;\n> >  \t}\n> >  \tobj->size = size;\n> > diff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\n> > index e01cf6e360..59a36c2481 100644\n> > --- a/builtin/unpack-objects.c\n> > +++ b/builtin/unpack-objects.c\n> > @@ -533,7 +533,8 @@ static void unpack_one(unsigned nr)\n> >  {\n> >  \tunsigned shift;\n> >  \tunsigned char *pack;\n> > -\tunsigned long size, c;\n> > +\tsize_t size;\n> > +\tunsigned long c;\n> >  \tenum object_type type;\n> >  \n> >  \tobj_list[nr].offset = consumed_bytes;\n> > @@ -548,7 +549,7 @@ static void unpack_one(unsigned nr)\n> >  \t\tpack = fill(1);\n> >  \t\tc = *pack;\n> >  \t\tuse(1);\n> > -\t\tsize += (c & 0x7f) << shift;\n> > +\t\tsize += ((size_t)c & 0x7f) << shift;\n> >  \t\tshift += 7;\n> >  \t}\n> >  \n> > -- \n> > gitgitgadget\n> > \n> > \n> \n"},{"id":"542888","messageId":"34a8b2ec-b651-18b3-c62d-072f49dab33e@gmx.de","threadId":"65563","inReplyTo":"20260505192722.GB12275@tb-raspi4","subject":"Re: [PATCH v2 03/11] odb, packfile: use size_t for streaming object sizes","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2026-05-08T07:38:43Z","receivedAt":"2026-05-08T07:38:47Z","isPatch":true,"body":"Hi Torsten,\n\nOn Tue, 5 May 2026, Torsten Bögershausen wrote:\n\n> On Mon, May 04, 2026 at 05:08:20PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > \n> > The odb_read_stream structure uses unsigned long for the size field,\n> > which is 32-bit on Windows even in 64-bit builds. When streaming\n> > objects larger than 4GB, the size would be truncated to zero or an\n> > incorrect value, resulting in empty files being written to disk.\n> > \n> > Change the size field in odb_read_stream to size_t and introduce\n> > unpack_object_header_sz() to return sizes via size_t pointer. Since\n> > object_info.sizep remains unsigned long for API compatibility, use\n> > temporary variables where the types differ, with comments noting the\n> > truncation limitation for code paths that still use unsigned long.\n> > \n> > This was originally authored by LordKiRon <https://github.com/LordKiRon>,\n> > who preferred not to reveal their real name and therefore agreed that I\n> > take over authorship.\n> > \n> > Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > ---\n> >  builtin/pack-objects.c       | 23 ++++++++++++++++-------\n> >  object-file.c                | 10 +++++++++-\n> >  odb/streaming.c              | 13 ++++++++++++-\n> >  odb/streaming.h              |  2 +-\n> >  oss-fuzz/fuzz-pack-headers.c |  2 +-\n> >  pack-bitmap.c                |  2 +-\n> >  pack-check.c                 |  6 ++++--\n> >  packfile.c                   | 24 +++++++++++++++---------\n> >  packfile.h                   |  4 ++--\n> >  9 files changed, 61 insertions(+), 25 deletions(-)\n> > \n> \n> > diff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\n> > index dd2480a73d..aa4b1cb9b8 100644\n> > --- a/builtin/pack-objects.c\n> > +++ b/builtin/pack-objects.c\n> \n> I haven't been able to follow all changes, so this may be false alarm.\n> Do we need a cast_size_t_to_ulong() somewhere ?\n\nOh, that's a really good catch, I forgot to add those. There are a couple\nof callers, indeed, and they still have an implicit, _silent_ truncation\nto `unsigned long` that I'd rather have explicit, with a check. I'll add\nthose callers in the next iteration.\n\nThank you!\nJohannes\n\n> \n> > @@ -629,14 +629,21 @@ static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n> >  \tstruct packed_git *p = IN_PACK(entry);\n> >  \tstruct pack_window *w_curs = NULL;\n> >  \tuint32_t pos;\n> > -\toff_t offset;\n> > +\toff_t offset, cur;\n> >  \tenum object_type type = oe_type(entry);\n> > +\tenum object_type in_pack_type;\n> >  \toff_t datalen;\n> >  \tunsigned char header[MAX_PACK_OBJECT_HEADER],\n> >  \t\t      dheader[MAX_PACK_OBJECT_HEADER];\n> >  \tunsigned hdrlen;\n> >  \tconst unsigned hashsz = the_hash_algo->rawsz;\n> > -\tunsigned long entry_size = SIZE(entry);\n> > +\tsize_t entry_size;\n> > +\n> > +\tcur = entry->in_pack_offset;\n> > +\tin_pack_type = unpack_object_header(p, &w_curs, &cur, &entry_size);\n> > +\tif (in_pack_type < 0)\n> > +\t\tdie(_(\"write_reuse_object: unable to parse object header of %s\"),\n> > +\t\t    oid_to_hex(&entry->idx.oid));\n> >  \n> >  \tif (DELTA(entry))\n> >  \t\ttype = (allow_ofs_delta && DELTA(entry)->idx.offset) ?\n> > @@ -1087,7 +1094,7 @@ static void write_reused_pack_one(struct packed_git *reuse_packfile,\n> >  {\n> >  \toff_t offset, next, cur;\n> >  \tenum object_type type;\n> > -\tunsigned long size;\n> > +\tsize_t size;\n> >  \n> >  \toffset = pack_pos_to_offset(reuse_packfile, pos);\n> >  \tnext = pack_pos_to_offset(reuse_packfile, pos + 1);\n> > @@ -2243,7 +2250,7 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n> >  \t\toff_t ofs;\n> >  \t\tunsigned char *buf, c;\n> >  \t\tenum object_type type;\n> > -\t\tunsigned long in_pack_size;\n> > +\t\tsize_t in_pack_size;\n> >  \n> >  \t\tbuf = use_pack(p, &w_curs, entry->in_pack_offset, &avail);\n> >  \n> > @@ -2734,16 +2741,18 @@ unsigned long oe_get_size_slow(struct packing_data *pack,\n> >  \tstruct pack_window *w_curs;\n> >  \tunsigned char *buf;\n> >  \tenum object_type type;\n> > -\tunsigned long used, avail, size;\n> > +\tunsigned long used, avail;\n> > +\tsize_t size;\n> >  \n> >  \tif (e->type_ != OBJ_OFS_DELTA && e->type_ != OBJ_REF_DELTA) {\n> > +\t\tunsigned long sz;\n> >  \t\tpacking_data_lock(&to_pack);\n> >  \t\tif (odb_read_object_info(the_repository->objects,\n> > -\t\t\t\t\t &e->idx.oid, &size) < 0)\n> > +\t\t\t\t\t &e->idx.oid, &sz) < 0)\n> >  \t\t\tdie(_(\"unable to get size of %s\"),\n> >  \t\t\t    oid_to_hex(&e->idx.oid));\n> >  \t\tpacking_data_unlock(&to_pack);\n> > -\t\treturn size;\n> > +\t\treturn sz;\n> >  \t}\n> >  \n> >  \tp = oe_in_pack(pack, e);\n> > diff --git a/object-file.c b/object-file.c\n> > index 086b2b65ff..0be2981c7a 100644\n> > --- a/object-file.c\n> > +++ b/object-file.c\n> > @@ -2326,6 +2326,7 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n> >  \tstruct object_info oi = OBJECT_INFO_INIT;\n> >  \tstruct odb_loose_read_stream *st;\n> >  \tunsigned long mapsize;\n> > +\tunsigned long size_ul;\n> >  \tvoid *mapped;\n> >  \n> >  \tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n> > @@ -2349,11 +2350,18 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n> >  \t\tgoto error;\n> >  \t}\n> >  \n> > -\toi.sizep = &st->base.size;\n> > +\t/*\n> > +\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n> > +\t * st->base.size is size_t (64-bit). Use temporary variable.\n> > +\t * Note: loose objects >4GB would still truncate here, but such\n> > +\t * large loose objects are uncommon (they'd normally be packed).\n> > +\t */\n> > +\toi.sizep = &size_ul;\n> >  \toi.typep = &st->base.type;\n> >  \n> >  \tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n> >  \t\tgoto error;\n> > +\tst->base.size = size_ul;\n> >  \n> >  \tst->mapped = mapped;\n> >  \tst->mapsize = mapsize;\n> > diff --git a/odb/streaming.c b/odb/streaming.c\n> > index 5927a12954..af2adf5ce7 100644\n> > --- a/odb/streaming.c\n> > +++ b/odb/streaming.c\n> > @@ -157,15 +157,26 @@ static int open_istream_incore(struct odb_read_stream **out,\n> >  \t\t.base.read = read_istream_incore,\n> >  \t};\n> >  \tstruct odb_incore_read_stream *st;\n> > +\tunsigned long size_ul;\n> >  \tint ret;\n> >  \n> >  \toi.typep = &stream.base.type;\n> > -\toi.sizep = &stream.base.size;\n> > +\t/*\n> > +\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n> > +\t * stream.base.size is size_t (64-bit). We use a temporary variable\n> > +\t * because the types are incompatible. Note: this path still truncates\n> > +\t * for >4GB objects, but large objects should use pack streaming\n> > +\t * (packfile_store_read_object_stream) which handles size_t properly.\n> > +\t * This incore fallback is only used for small objects or when pack\n> > +\t * streaming is unavailable.\n> > +\t */\n> > +\toi.sizep = &size_ul;\n> >  \toi.contentp = (void **)&stream.buf;\n> >  \tret = odb_read_object_info_extended(odb, oid, &oi,\n> >  \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n> >  \tif (ret)\n> >  \t\treturn ret;\n> > +\tstream.base.size = size_ul;\n> >  \n> >  \tCALLOC_ARRAY(st, 1);\n> >  \t*st = stream;\n> > diff --git a/odb/streaming.h b/odb/streaming.h\n> > index c7861f7e13..517e2ea2d3 100644\n> > --- a/odb/streaming.h\n> > +++ b/odb/streaming.h\n> > @@ -21,7 +21,7 @@ struct odb_read_stream {\n> >  \todb_read_stream_close_fn close;\n> >  \todb_read_stream_read_fn read;\n> >  \tenum object_type type;\n> > -\tunsigned long size; /* inflated size of full object */\n> > +\tsize_t size; /* inflated size of full object */\n> >  };\n> >  \n> >  /*\n> > diff --git a/oss-fuzz/fuzz-pack-headers.c b/oss-fuzz/fuzz-pack-headers.c\n> > index 150c0f5fa2..ef61ab577c 100644\n> > --- a/oss-fuzz/fuzz-pack-headers.c\n> > +++ b/oss-fuzz/fuzz-pack-headers.c\n> > @@ -6,7 +6,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size);\n> >  int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n> >  {\n> >  \tenum object_type type;\n> > -\tunsigned long len;\n> > +\tsize_t len;\n> >  \n> >  \tunpack_object_header_buffer((const unsigned char *)data,\n> >  \t\t\t\t    (unsigned long)size, &type, &len);\n> > diff --git a/pack-bitmap.c b/pack-bitmap.c\n> > index f6ec18d83a..f9af8a96bd 100644\n> > --- a/pack-bitmap.c\n> > +++ b/pack-bitmap.c\n> > @@ -2270,7 +2270,7 @@ static int try_partial_reuse(struct bitmap_index *bitmap_git,\n> >  {\n> >  \toff_t delta_obj_offset;\n> >  \tenum object_type type;\n> > -\tunsigned long size;\n> > +\tsize_t size;\n> >  \n> >  \tif (pack_pos >= pack->p->num_objects)\n> >  \t\treturn -1; /* not actually in the pack */\n> > diff --git a/pack-check.c b/pack-check.c\n> > index 79992bb509..2792f34d25 100644\n> > --- a/pack-check.c\n> > +++ b/pack-check.c\n> > @@ -110,7 +110,7 @@ static int verify_packfile(struct repository *r,\n> >  \t\tvoid *data;\n> >  \t\tstruct object_id oid;\n> >  \t\tenum object_type type;\n> > -\t\tunsigned long size;\n> > +\t\tsize_t size;\n> >  \t\toff_t curpos;\n> >  \t\tint data_valid;\n> >  \n> > @@ -143,7 +143,9 @@ static int verify_packfile(struct repository *r,\n> >  \t\t\tdata = NULL;\n> >  \t\t\tdata_valid = 0;\n> >  \t\t} else {\n> > -\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &size);\n> > +\t\t\tunsigned long sz;\n> > +\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &sz);\n> > +\t\t\tsize = sz;\n> >  \t\t\tdata_valid = 1;\n> >  \t\t}\n> >  \n> > diff --git a/packfile.c b/packfile.c\n> > index b012d648ad..fdae91dd11 100644\n> > --- a/packfile.c\n> > +++ b/packfile.c\n> > @@ -1133,7 +1133,7 @@ out:\n> >  }\n> >  \n> >  unsigned long unpack_object_header_buffer(const unsigned char *buf,\n> > -\t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n> > +\t\tunsigned long len, enum object_type *type, size_t *sizep)\n> >  {\n> >  \tunsigned shift;\n> >  \tsize_t size, c;\n> > @@ -1144,7 +1144,11 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n> >  \tsize = c & 15;\n> >  \tshift = 4;\n> >  \twhile (c & 0x80) {\n> > -\t\tif (len <= used || (bitsizeof(long) - 7) < shift) {\n> > +\t\t/*\n> > +\t\t * Each continuation byte adds 7 bits. Ensure shift won't\n> > +\t\t * overflow size_t (use size_t not long for 64-bit on Windows).\n> > +\t\t */\n> > +\t\tif (len <= used || (bitsizeof(size_t) - 7) < shift) {\n> >  \t\t\terror(\"bad object header\");\n> >  \t\t\tsize = used = 0;\n> >  \t\t\tbreak;\n> > @@ -1153,7 +1157,7 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n> >  \t\tsize = st_add(size, st_left_shift(c & 0x7f, shift));\n> >  \t\tshift += 7;\n> >  \t}\n> > -\t*sizep = cast_size_t_to_ulong(size);\n> > +\t*sizep = size;\n> >  \treturn used;\n> >  }\n> >  \n> > @@ -1215,7 +1219,7 @@ unsigned long get_size_from_delta(struct packed_git *p,\n> >  int unpack_object_header(struct packed_git *p,\n> >  \t\t\t struct pack_window **w_curs,\n> >  \t\t\t off_t *curpos,\n> > -\t\t\t unsigned long *sizep)\n> > +\t\t\t size_t *sizep)\n> >  {\n> >  \tunsigned char *base;\n> >  \tunsigned long left;\n> > @@ -1367,7 +1371,7 @@ static enum object_type packed_to_object_type(struct repository *r,\n> >  \n> >  \twhile (type == OBJ_OFS_DELTA || type == OBJ_REF_DELTA) {\n> >  \t\toff_t base_offset;\n> > -\t\tunsigned long size;\n> > +\t\tsize_t size;\n> >  \t\t/* Push the object we're going to leave behind */\n> >  \t\tif (poi_stack_nr >= poi_stack_alloc && poi_stack == small_poi_stack) {\n> >  \t\t\tpoi_stack_alloc = alloc_nr(poi_stack_nr);\n> > @@ -1586,7 +1590,7 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n> >  \t\t\t\t\t     uint32_t *maybe_index_pos, struct object_info *oi)\n> >  {\n> >  \tstruct pack_window *w_curs = NULL;\n> > -\tunsigned long size;\n> > +\tsize_t size;\n> >  \toff_t curpos = obj_offset;\n> >  \tenum object_type type = OBJ_NONE;\n> >  \tuint32_t pack_pos;\n> > @@ -1778,7 +1782,7 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n> >  \tstruct pack_window *w_curs = NULL;\n> >  \toff_t curpos = obj_offset;\n> >  \tvoid *data = NULL;\n> > -\tunsigned long size;\n> > +\tsize_t size;\n> >  \tenum object_type type;\n> >  \tstruct unpack_entry_stack_ent small_delta_stack[UNPACK_ENTRY_STACK_PREALLOC];\n> >  \tstruct unpack_entry_stack_ent *delta_stack = small_delta_stack;\n> > @@ -1943,8 +1947,10 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n> >  \t\t\t      (uintmax_t)curpos, p->pack_name);\n> >  \t\t\tdata = NULL;\n> >  \t\t} else {\n> > +\t\t\tunsigned long sz;\n> >  \t\t\tdata = patch_delta(base, base_size, delta_data,\n> > -\t\t\t\t\t   delta_size, &size);\n> > +\t\t\t\t\t   delta_size, &sz);\n> > +\t\t\tsize = sz;\n> >  \n> >  \t\t\t/*\n> >  \t\t\t * We could not apply the delta; warn the user, but\n> > @@ -2929,7 +2935,7 @@ int packfile_read_object_stream(struct odb_read_stream **out,\n> >  \tstruct odb_packed_read_stream *stream;\n> >  \tstruct pack_window *window = NULL;\n> >  \tenum object_type in_pack_type;\n> > -\tunsigned long size;\n> > +\tsize_t size;\n> >  \n> >  \tin_pack_type = unpack_object_header(pack, &window, &offset, &size);\n> >  \tunuse_pack(&window);\n> > diff --git a/packfile.h b/packfile.h\n> > index 9b647da7dd..49d6bdecf6 100644\n> > --- a/packfile.h\n> > +++ b/packfile.h\n> > @@ -456,9 +456,9 @@ off_t find_pack_entry_one(const struct object_id *oid, struct packed_git *);\n> >  \n> >  int is_pack_valid(struct packed_git *);\n> >  void *unpack_entry(struct repository *r, struct packed_git *, off_t, enum object_type *, unsigned long *);\n> > -unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n> > +unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, size_t *sizep);\n> >  unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n> > -int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\n> > +int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, size_t *);\n> >  off_t get_delta_base(struct packed_git *p, struct pack_window **w_curs,\n> >  \t\t     off_t *curpos, enum object_type type,\n> >  \t\t     off_t delta_obj_offset);\n> > -- \n> > gitgitgadget\n> > \n> > \n> \n"},{"id":"542889","messageId":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v2.git.1777914508.gitgitgadget@gmail.com","subject":"[PATCH v3 00/11] Handle cloning of objects larger than 4GB on Windows","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:38Z","receivedAt":"2026-05-08T08:16:53Z","isPatch":true,"body":"On Windows, unsigned long is 32-bit even on 64-bit systems. This causes\nmultiple problems when Git handles objects larger than 4GB. This patch\nseries is a very targeted fix for a very early part of the problem: it\naddresses the most fundamental truncation points that prevent a >4GB object\nfrom surviving a clone at all.\n\nSpecifically, this fixes:\n\n * zlib's uLong wrapping and triggering BUG() assertions in the git_zstream\n   wrapper\n * Object sizes being truncated in pack streaming, delta headers, and\n   index-pack/unpack-objects\n * pack-objects re-encoding reused pack entries with a truncated size,\n   producing corrupt packs on the wire\n\nMany other code paths still use unsigned long for object sizes (e.g.,\ncat-file -s, object_info.sizep, the delta machinery) and will need their own\nconversions. This series does not attempt to fix those.\n\nBased on work by @LordKiRon in git-for-windows/git#6076.\n\nFor testing, add a test helper that synthesizes a pack with a >4GB blob and\nregression tests that clone it via both the unpack-objects and index-pack\ncode paths using file:// transport. Since these test cases are quite slow\n(even after optimizing the pack generation part, the git clone test has no\nchance but to hash 2x4GB of data), they are marked as EXPENSIVE. To ensure\nthat they are passing well in advance of any release, the CI is changed to\nrun them in the CI builds of relatively infrequent integration branch\nupdates.\n\nChanges since v2:\n\n * Now uses the proper data type for the varint decoding value (thanks,\n   Torsten!)\n * The callers that now would silently narrow size_t to unsigned long\n   properly check and error out instead (thanks, Torsten!)\n\nChanges since v1:\n\n * dramatically accelerated the test helper that generates 4GB pack files,\n   via two separate strategies:\n   1. using the \"unsafe\" SHA-1 for the blob OID computation.\n   2. using pre-computed \"Lego blocks\" to construct the 4GB packs needed in\n      the test cases, where the size (and therefore the involved OIDs) are\n      well-known in advance.\n * even with these improvements, the actual git clone is still slow (of\n   course, because it cannot use any of those shortcuts), therefore the\n   tests are marked as EXPENSIVE.\n * to exercise those tests nevertheless, the last patch lets all EXPENSIVE\n   test cases be run for the integration branches other than seen.\n\nJohannes Schindelin (11):\n  index-pack, unpack-objects: use size_t for object size\n  git-zlib: handle data streams larger than 4GB\n  odb, packfile: use size_t for streaming object sizes\n  delta, packfile: use size_t for delta header sizes\n  test-tool: add a helper to synthesize large packfiles\n  t5608: add regression test for >4GB object clone\n  test-tool synthesize: use the unsafe hash for speed\n  test-tool synthesize: precompute pack for 4 GiB + 1\n  test-tool synthesize: add precomputed SHA-256 pack for 4 GiB + 1\n  t5608: mark >4GB tests as EXPENSIVE\n  ci: run expensive tests on push builds to integration branches\n\n Makefile                     |   1 +\n builtin/index-pack.c         |   8 +-\n builtin/pack-objects.c       |  34 ++-\n builtin/unpack-objects.c     |   4 +-\n ci/lib.sh                    |   9 +\n compat/zlib-compat.h         |   2 +\n delta.h                      |  14 +-\n git-zlib.c                   |  25 +-\n git-zlib.h                   |   4 +-\n object-file.c                |  12 +-\n odb/streaming.c              |  13 +-\n odb/streaming.h              |   2 +-\n oss-fuzz/fuzz-pack-headers.c |   2 +-\n pack-bitmap.c                |   2 +-\n pack-check.c                 |   6 +-\n packfile.c                   |  57 ++--\n packfile.h                   |   4 +-\n t/helper/meson.build         |   1 +\n t/helper/test-synthesize.c   | 541 +++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c         |   1 +\n t/helper/test-tool.h         |   1 +\n t/t5608-clone-2gb.sh         |  37 +++\n 22 files changed, 724 insertions(+), 56 deletions(-)\n create mode 100644 t/helper/test-synthesize.c\n\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2102%2Fdscho%2Ffix-large-clones-on-windows-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2102/dscho/fix-large-clones-on-windows-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/2102\n\nRange-diff vs v2:\n\n  1:  dc660106ea !  1:  311cdc601d index-pack, unpack-objects: use size_t for object size\n     @@ Commit message\n          64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\n          space.\n      \n     +    Declare the per-byte continuation variable `c` as size_t as well,\n     +    matching the canonical varint decoder unpack_object_header_buffer()\n     +    in packfile.c. With c as size_t the expression (c & 0x7f) << shift\n     +    is naturally size_t-typed, so the explicit cast that an earlier\n     +    iteration carried at the use site is no longer needed.\n     +\n     +    While at it, add the same overflow guard that\n     +    unpack_object_header_buffer() carries: if the cumulative shift would\n     +    exceed bitsizeof(size_t) - 7, refuse the input rather than invoking\n     +    undefined behavior. Unlike unpack_object_header_buffer(), which\n     +    labels this case \"bad object header\", report it as the platform\n     +    limit it actually is: a header may be perfectly well-formed and\n     +    still encode a size we cannot represent locally (notably on a\n     +    32-bit build consuming a packfile produced on a 64-bit host).\n     +\n          This was originally authored by LordKiRon <https://github.com/LordKiRon>,\n          who preferred not to reveal their real name and therefore agreed that I\n          take over authorship.\n      \n     +    Helped-by: Torsten Bögershausen <tboegi@web.de>\n          Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## builtin/index-pack.c ##\n     @@ builtin/index-pack.c: static void *unpack_raw_entry(struct object_entry *obj,\n       {\n       \tunsigned char *p;\n      -\tunsigned long size, c;\n     -+\tsize_t size;\n     -+\tunsigned long c;\n     ++\tsize_t size, c;\n       \toff_t base_offset;\n       \tunsigned shift;\n       \tvoid *data;\n      @@ builtin/index-pack.c: static void *unpack_raw_entry(struct object_entry *obj,\n     + \tsize = (c & 15);\n     + \tshift = 4;\n     + \twhile (c & 0x80) {\n     ++\t\tif ((bitsizeof(size_t) - 7) < shift)\n     ++\t\t\tdie(_(\"object size too large for this platform\"));\n       \t\tp = fill(1);\n       \t\tc = *p;\n       \t\tuse(1);\n     --\t\tsize += (c & 0x7f) << shift;\n     -+\t\tsize += ((size_t)c & 0x7f) << shift;\n     - \t\tshift += 7;\n     - \t}\n     - \tobj->size = size;\n      \n       ## builtin/unpack-objects.c ##\n      @@ builtin/unpack-objects.c: static void unpack_one(unsigned nr)\n     @@ builtin/unpack-objects.c: static void unpack_one(unsigned nr)\n       \tunsigned shift;\n       \tunsigned char *pack;\n      -\tunsigned long size, c;\n     -+\tsize_t size;\n     -+\tunsigned long c;\n     ++\tsize_t size, c;\n       \tenum object_type type;\n       \n       \tobj_list[nr].offset = consumed_bytes;\n      @@ builtin/unpack-objects.c: static void unpack_one(unsigned nr)\n     + \tsize = (c & 15);\n     + \tshift = 4;\n     + \twhile (c & 0x80) {\n     ++\t\tif ((bitsizeof(size_t) - 7) < shift)\n     ++\t\t\tdie(_(\"object size too large for this platform\"));\n       \t\tpack = fill(1);\n       \t\tc = *pack;\n       \t\tuse(1);\n     --\t\tsize += (c & 0x7f) << shift;\n     -+\t\tsize += ((size_t)c & 0x7f) << shift;\n     - \t\tshift += 7;\n     - \t}\n     - \n  2:  92f4327b1f =  2:  c611913194 git-zlib: handle data streams larger than 4GB\n  3:  3a539061c5 !  3:  b789f57de9 odb, packfile: use size_t for streaming object sizes\n     @@ Commit message\n          temporary variables where the types differ, with comments noting the\n          truncation limitation for code paths that still use unsigned long.\n      \n     +    Widening the producers to size_t in this way introduces a handful of\n     +    silent size_t -> unsigned long narrowings on Windows, all in\n     +    builtin/pack-objects.c, where the consumers are still typed\n     +    unsigned long. Make those narrowings explicit with\n     +    cast_size_t_to_ulong() so they assert loudly the moment an object\n     +    actually exceeds ULONG_MAX bytes:\n     +\n     +      - oe_get_size_slow() returns unsigned long but holds a size_t\n     +        locally; cast at the return.\n     +      - write_reuse_object() passes a size_t into check_pack_inflate(),\n     +        whose expect parameter is unsigned long; cast at the call.\n     +      - check_object() routes a size_t through SET_SIZE() and\n     +        SET_DELTA_SIZE(), both of which take unsigned long via\n     +        oe_set_size() / oe_set_delta_size(); cast at the three call\n     +        sites in the OBJ_OFS_DELTA / OBJ_REF_DELTA branches and in the\n     +        non-delta default arm.\n     +\n     +    The cast-only treatment is deliberately a stop-gap. Properly\n     +    widening oe_set_size, oe_get_size_slow's return type,\n     +    check_pack_inflate's expect parameter, object_info.sizep,\n     +    patch_delta, and the OE_SIZE_BITS bit-fields cascades into a series\n     +    that is too large to be reviewable, so the proper widening is\n     +    deferred to a follow-up topic. Until then,\n     +    cast_size_t_to_ulong() at least makes the truncation explicit at\n     +    the source: it documents the boundary, and on a 64-bit non-Windows\n     +    platform it is a no-op.\n     +\n          This was originally authored by LordKiRon <https://github.com/LordKiRon>,\n          who preferred not to reveal their real name and therefore agreed that I\n          take over authorship.\n      \n     +    Helped-by: Torsten Bögershausen <tboegi@web.de>\n          Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n      \n       ## builtin/pack-objects.c ##\n     @@ builtin/pack-objects.c: static off_t write_reuse_object(struct hashfile *f, stru\n       \n       \tif (DELTA(entry))\n       \t\ttype = (allow_ofs_delta && DELTA(entry)->idx.offset) ?\n     +@@ builtin/pack-objects.c: static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n     + \tdatalen -= entry->in_pack_header_size;\n     + \n     + \tif (!pack_to_stdout && p->index_version == 1 &&\n     +-\t    check_pack_inflate(p, &w_curs, offset, datalen, entry_size)) {\n     ++\t    check_pack_inflate(p, &w_curs, offset, datalen,\n     ++\t\t\t       cast_size_t_to_ulong(entry_size))) {\n     + \t\terror(_(\"corrupt packed object for %s\"),\n     + \t\t      oid_to_hex(&entry->idx.oid));\n     + \t\tunuse_pack(&w_curs);\n      @@ builtin/pack-objects.c: static void write_reused_pack_one(struct packed_git *reuse_packfile,\n       {\n       \toff_t offset, next, cur;\n     @@ builtin/pack-objects.c: static void check_object(struct object_entry *entry, uin\n       \n       \t\tbuf = use_pack(p, &w_curs, entry->in_pack_offset, &avail);\n       \n     +@@ builtin/pack-objects.c: static void check_object(struct object_entry *entry, uint32_t object_index)\n     + \t\tdefault:\n     + \t\t\t/* Not a delta hence we've already got all we need. */\n     + \t\t\toe_set_type(entry, entry->in_pack_type);\n     +-\t\t\tSET_SIZE(entry, in_pack_size);\n     ++\t\t\tSET_SIZE(entry, cast_size_t_to_ulong(in_pack_size));\n     + \t\t\tentry->in_pack_header_size = used;\n     + \t\t\tif (oe_type(entry) < OBJ_COMMIT || oe_type(entry) > OBJ_BLOB)\n     + \t\t\t\tgoto give_up;\n     +@@ builtin/pack-objects.c: static void check_object(struct object_entry *entry, uint32_t object_index)\n     + \t\tif (have_base &&\n     + \t\t    can_reuse_delta(&base_ref, entry, &base_entry)) {\n     + \t\t\toe_set_type(entry, entry->in_pack_type);\n     +-\t\t\tSET_SIZE(entry, in_pack_size); /* delta size */\n     +-\t\t\tSET_DELTA_SIZE(entry, in_pack_size);\n     ++\t\t\tSET_SIZE(entry, cast_size_t_to_ulong(in_pack_size)); /* delta size */\n     ++\t\t\tSET_DELTA_SIZE(entry, cast_size_t_to_ulong(in_pack_size));\n     + \n     + \t\t\tif (base_entry) {\n     + \t\t\t\tSET_DELTA(entry, base_entry);\n      @@ builtin/pack-objects.c: unsigned long oe_get_size_slow(struct packing_data *pack,\n       \tstruct pack_window *w_curs;\n       \tunsigned char *buf;\n     @@ builtin/pack-objects.c: unsigned long oe_get_size_slow(struct packing_data *pack\n       \t}\n       \n       \tp = oe_in_pack(pack, e);\n     +@@ builtin/pack-objects.c: unsigned long oe_get_size_slow(struct packing_data *pack,\n     + \n     + \tunuse_pack(&w_curs);\n     + \tpacking_data_unlock(&to_pack);\n     +-\treturn size;\n     ++\treturn cast_size_t_to_ulong(size);\n     + }\n     + \n     + static int try_delta(struct unpacked *trg, struct unpacked *src,\n      \n       ## object-file.c ##\n      @@ object-file.c: int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n  4:  3274cba862 =  4:  8e87a4e71f delta, packfile: use size_t for delta header sizes\n  5:  afa74a3a2b =  5:  34fec4a32d test-tool: add a helper to synthesize large packfiles\n  6:  a3019888d8 =  6:  88f992903f t5608: add regression test for >4GB object clone\n  7:  859e93e7a9 =  7:  4f207c8a47 test-tool synthesize: use the unsafe hash for speed\n  8:  29b9a74e91 =  8:  2751c21c6e test-tool synthesize: precompute pack for 4 GiB + 1\n  9:  8e6e720804 =  9:  3a006d96c3 test-tool synthesize: add precomputed SHA-256 pack for 4 GiB + 1\n 10:  5b44410b2f = 10:  86c09af4f5 t5608: mark >4GB tests as EXPENSIVE\n 11:  1eaaa7fad7 = 11:  2159f6a271 ci: run expensive tests on push builds to integration branches\n\n-- \ngitgitgadget\n"},{"id":"542890","messageId":"311cdc601d089bc96e17dc008780a1e7e54fa49d.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:39Z","receivedAt":"2026-05-08T08:16:54Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nWhen unpacking objects from a packfile, the object size is decoded\nfrom a variable-length encoding. On platforms where unsigned long is\n32-bit (such as Windows, even in 64-bit builds), the shift operation\noverflows when decoding sizes larger than 4GB. The result is a\ntruncated size value, causing the unpacked object to be corrupted or\nrejected.\n\nFix this by changing the size variable to size_t, which is 64-bit on\n64-bit platforms, and ensuring the shift arithmetic occurs in 64-bit\nspace.\n\nDeclare the per-byte continuation variable `c` as size_t as well,\nmatching the canonical varint decoder unpack_object_header_buffer()\nin packfile.c. With c as size_t the expression (c & 0x7f) << shift\nis naturally size_t-typed, so the explicit cast that an earlier\niteration carried at the use site is no longer needed.\n\nWhile at it, add the same overflow guard that\nunpack_object_header_buffer() carries: if the cumulative shift would\nexceed bitsizeof(size_t) - 7, refuse the input rather than invoking\nundefined behavior. Unlike unpack_object_header_buffer(), which\nlabels this case \"bad object header\", report it as the platform\nlimit it actually is: a header may be perfectly well-formed and\nstill encode a size we cannot represent locally (notably on a\n32-bit build consuming a packfile produced on a 64-bit host).\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nHelped-by: Torsten Bögershausen <tboegi@web.de>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n builtin/index-pack.c     | 8 +++++---\n builtin/unpack-objects.c | 4 +++-\n 2 files changed, 8 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..2e4b42fa12 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -37,7 +37,7 @@ static const char index_pack_usage[] =\n \n struct object_entry {\n \tstruct pack_idx_entry idx;\n-\tunsigned long size;\n+\tsize_t size;\n \tunsigned char hdr_size;\n \tsigned char type;\n \tsigned char real_type;\n@@ -469,7 +469,7 @@ static int is_delta_type(enum object_type type)\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n-static void *unpack_entry_data(off_t offset, unsigned long size,\n+static void *unpack_entry_data(off_t offset, size_t size,\n \t\t\t       enum object_type type, struct object_id *oid)\n {\n \tstatic char fixed_buf[8192];\n@@ -524,7 +524,7 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \t\t\t      struct object_id *oid)\n {\n \tunsigned char *p;\n-\tunsigned long size, c;\n+\tsize_t size, c;\n \toff_t base_offset;\n \tunsigned shift;\n \tvoid *data;\n@@ -539,6 +539,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \tsize = (c & 15);\n \tshift = 4;\n \twhile (c & 0x80) {\n+\t\tif ((bitsizeof(size_t) - 7) < shift)\n+\t\t\tdie(_(\"object size too large for this platform\"));\n \t\tp = fill(1);\n \t\tc = *p;\n \t\tuse(1);\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..76b3d0dee3 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -533,7 +533,7 @@ static void unpack_one(unsigned nr)\n {\n \tunsigned shift;\n \tunsigned char *pack;\n-\tunsigned long size, c;\n+\tsize_t size, c;\n \tenum object_type type;\n \n \tobj_list[nr].offset = consumed_bytes;\n@@ -545,6 +545,8 @@ static void unpack_one(unsigned nr)\n \tsize = (c & 15);\n \tshift = 4;\n \twhile (c & 0x80) {\n+\t\tif ((bitsizeof(size_t) - 7) < shift)\n+\t\t\tdie(_(\"object size too large for this platform\"));\n \t\tpack = fill(1);\n \t\tc = *pack;\n \t\tuse(1);\n-- \ngitgitgadget\n\n"},{"id":"542891","messageId":"c611913194cab1fcba5f990bf44ba15f721e1223.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 02/11] git-zlib: handle data streams larger than 4GB","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:40Z","receivedAt":"2026-05-08T08:16:56Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nOn Windows, zlib's `uLong` type is 32-bit even on 64-bit systems. When\nprocessing data streams larger than 4GB, the `total_in` and `total_out`\nfields in zlib's `z_stream` structure wrap around, which caused the\nsanity checks in `zlib_post_call()` to trigger `BUG()` assertions.\n\nThe git_zstream wrapper now tracks its own 64-bit totals rather than\ncopying them from zlib. The sanity checks compare only the low bits,\nusing `maximum_unsigned_value_of_type(uLong)` to mask appropriately for\nthe platform's `uLong` size.\n\nThis is based on work by LordKiRon in git-for-windows#6076.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n git-zlib.c    | 25 +++++++++++++++++--------\n git-zlib.h    |  4 ++--\n object-file.c |  2 +-\n 3 files changed, 20 insertions(+), 11 deletions(-)\n\ndiff --git a/git-zlib.c b/git-zlib.c\nindex df9604910e..b91cb323ae 100644\n--- a/git-zlib.c\n+++ b/git-zlib.c\n@@ -30,6 +30,9 @@ static const char *zerr_to_string(int status)\n  */\n /* #define ZLIB_BUF_MAX ((uInt)-1) */\n #define ZLIB_BUF_MAX ((uInt) 1024 * 1024 * 1024) /* 1GB */\n+\n+/* uLong is 32-bit on Windows, even on 64-bit systems */\n+#define ULONG_MAX_VALUE maximum_unsigned_value_of_type(uLong)\n static inline uInt zlib_buf_cap(unsigned long len)\n {\n \treturn (ZLIB_BUF_MAX < len) ? ZLIB_BUF_MAX : len;\n@@ -39,31 +42,37 @@ static void zlib_pre_call(git_zstream *s)\n {\n \ts->z.next_in = s->next_in;\n \ts->z.next_out = s->next_out;\n-\ts->z.total_in = s->total_in;\n-\ts->z.total_out = s->total_out;\n+\ts->z.total_in = (uLong)(s->total_in & ULONG_MAX_VALUE);\n+\ts->z.total_out = (uLong)(s->total_out & ULONG_MAX_VALUE);\n \ts->z.avail_in = zlib_buf_cap(s->avail_in);\n \ts->z.avail_out = zlib_buf_cap(s->avail_out);\n }\n \n static void zlib_post_call(git_zstream *s, int status)\n {\n-\tunsigned long bytes_consumed;\n-\tunsigned long bytes_produced;\n+\tsize_t bytes_consumed;\n+\tsize_t bytes_produced;\n \n \tbytes_consumed = s->z.next_in - s->next_in;\n \tbytes_produced = s->z.next_out - s->next_out;\n-\tif (s->z.total_out != s->total_out + bytes_produced)\n+\t/*\n+\t * zlib's total_out/total_in are uLong which may wrap for >4GB.\n+\t * We track our own totals and verify only the low bits match.\n+\t */\n+\tif ((s->z.total_out & ULONG_MAX_VALUE) !=\n+\t    ((s->total_out + bytes_produced) & ULONG_MAX_VALUE))\n \t\tBUG(\"total_out mismatch\");\n \t/*\n \t * zlib does not update total_in when it returns Z_NEED_DICT,\n \t * causing a mismatch here. Skip the sanity check in that case.\n \t */\n \tif (status != Z_NEED_DICT &&\n-\t    s->z.total_in != s->total_in + bytes_consumed)\n+\t    (s->z.total_in & ULONG_MAX_VALUE) !=\n+\t    ((s->total_in + bytes_consumed) & ULONG_MAX_VALUE))\n \t\tBUG(\"total_in mismatch\");\n \n-\ts->total_out = s->z.total_out;\n-\ts->total_in = s->z.total_in;\n+\ts->total_out += bytes_produced;\n+\ts->total_in += bytes_consumed;\n \t/* zlib-ng marks `next_in` as `const`, so we have to cast it away. */\n \ts->next_in = (unsigned char *) s->z.next_in;\n \ts->next_out = s->z.next_out;\ndiff --git a/git-zlib.h b/git-zlib.h\nindex 0e66fefa8c..44380e8ad3 100644\n--- a/git-zlib.h\n+++ b/git-zlib.h\n@@ -7,8 +7,8 @@ typedef struct git_zstream {\n \tstruct z_stream_s z;\n \tunsigned long avail_in;\n \tunsigned long avail_out;\n-\tunsigned long total_in;\n-\tunsigned long total_out;\n+\tsize_t total_in;\n+\tsize_t total_out;\n \tunsigned char *next_in;\n \tunsigned char *next_out;\n } git_zstream;\ndiff --git a/object-file.c b/object-file.c\nindex 2acc9522df..086b2b65ff 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1118,7 +1118,7 @@ int odb_source_loose_write_stream(struct odb_source *source,\n \t} while (ret == Z_OK || ret == Z_BUF_ERROR);\n \n \tif (stream.total_in != len + hdrlen)\n-\t\tdie(_(\"write stream object %ld != %\"PRIuMAX), stream.total_in,\n+\t\tdie(_(\"write stream object %\"PRIuMAX\" != %\"PRIuMAX), (uintmax_t)stream.total_in,\n \t\t    (uintmax_t)len + hdrlen);\n \n \t/*\n-- \ngitgitgadget\n\n"},{"id":"542892","messageId":"b789f57de9dc21636db6dc6dd94ffebc2bbb351a.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 03/11] odb, packfile: use size_t for streaming object sizes","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:41Z","receivedAt":"2026-05-08T08:16:57Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe odb_read_stream structure uses unsigned long for the size field,\nwhich is 32-bit on Windows even in 64-bit builds. When streaming\nobjects larger than 4GB, the size would be truncated to zero or an\nincorrect value, resulting in empty files being written to disk.\n\nChange the size field in odb_read_stream to size_t and introduce\nunpack_object_header_sz() to return sizes via size_t pointer. Since\nobject_info.sizep remains unsigned long for API compatibility, use\ntemporary variables where the types differ, with comments noting the\ntruncation limitation for code paths that still use unsigned long.\n\nWidening the producers to size_t in this way introduces a handful of\nsilent size_t -> unsigned long narrowings on Windows, all in\nbuiltin/pack-objects.c, where the consumers are still typed\nunsigned long. Make those narrowings explicit with\ncast_size_t_to_ulong() so they assert loudly the moment an object\nactually exceeds ULONG_MAX bytes:\n\n  - oe_get_size_slow() returns unsigned long but holds a size_t\n    locally; cast at the return.\n  - write_reuse_object() passes a size_t into check_pack_inflate(),\n    whose expect parameter is unsigned long; cast at the call.\n  - check_object() routes a size_t through SET_SIZE() and\n    SET_DELTA_SIZE(), both of which take unsigned long via\n    oe_set_size() / oe_set_delta_size(); cast at the three call\n    sites in the OBJ_OFS_DELTA / OBJ_REF_DELTA branches and in the\n    non-delta default arm.\n\nThe cast-only treatment is deliberately a stop-gap. Properly\nwidening oe_set_size, oe_get_size_slow's return type,\ncheck_pack_inflate's expect parameter, object_info.sizep,\npatch_delta, and the OE_SIZE_BITS bit-fields cascades into a series\nthat is too large to be reviewable, so the proper widening is\ndeferred to a follow-up topic. Until then,\ncast_size_t_to_ulong() at least makes the truncation explicit at\nthe source: it documents the boundary, and on a 64-bit non-Windows\nplatform it is a no-op.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nHelped-by: Torsten Bögershausen <tboegi@web.de>\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n builtin/pack-objects.c       | 34 ++++++++++++++++++++++------------\n object-file.c                | 10 +++++++++-\n odb/streaming.c              | 13 ++++++++++++-\n odb/streaming.h              |  2 +-\n oss-fuzz/fuzz-pack-headers.c |  2 +-\n pack-bitmap.c                |  2 +-\n pack-check.c                 |  6 ++++--\n packfile.c                   | 24 +++++++++++++++---------\n packfile.h                   |  4 ++--\n 9 files changed, 67 insertions(+), 30 deletions(-)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex dd2480a73d..480cc0bd8c 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -629,14 +629,21 @@ static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n \tstruct packed_git *p = IN_PACK(entry);\n \tstruct pack_window *w_curs = NULL;\n \tuint32_t pos;\n-\toff_t offset;\n+\toff_t offset, cur;\n \tenum object_type type = oe_type(entry);\n+\tenum object_type in_pack_type;\n \toff_t datalen;\n \tunsigned char header[MAX_PACK_OBJECT_HEADER],\n \t\t      dheader[MAX_PACK_OBJECT_HEADER];\n \tunsigned hdrlen;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n-\tunsigned long entry_size = SIZE(entry);\n+\tsize_t entry_size;\n+\n+\tcur = entry->in_pack_offset;\n+\tin_pack_type = unpack_object_header(p, &w_curs, &cur, &entry_size);\n+\tif (in_pack_type < 0)\n+\t\tdie(_(\"write_reuse_object: unable to parse object header of %s\"),\n+\t\t    oid_to_hex(&entry->idx.oid));\n \n \tif (DELTA(entry))\n \t\ttype = (allow_ofs_delta && DELTA(entry)->idx.offset) ?\n@@ -664,7 +671,8 @@ static off_t write_reuse_object(struct hashfile *f, struct object_entry *entry,\n \tdatalen -= entry->in_pack_header_size;\n \n \tif (!pack_to_stdout && p->index_version == 1 &&\n-\t    check_pack_inflate(p, &w_curs, offset, datalen, entry_size)) {\n+\t    check_pack_inflate(p, &w_curs, offset, datalen,\n+\t\t\t       cast_size_t_to_ulong(entry_size))) {\n \t\terror(_(\"corrupt packed object for %s\"),\n \t\t      oid_to_hex(&entry->idx.oid));\n \t\tunuse_pack(&w_curs);\n@@ -1087,7 +1095,7 @@ static void write_reused_pack_one(struct packed_git *reuse_packfile,\n {\n \toff_t offset, next, cur;\n \tenum object_type type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \toffset = pack_pos_to_offset(reuse_packfile, pos);\n \tnext = pack_pos_to_offset(reuse_packfile, pos + 1);\n@@ -2243,7 +2251,7 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n \t\toff_t ofs;\n \t\tunsigned char *buf, c;\n \t\tenum object_type type;\n-\t\tunsigned long in_pack_size;\n+\t\tsize_t in_pack_size;\n \n \t\tbuf = use_pack(p, &w_curs, entry->in_pack_offset, &avail);\n \n@@ -2270,7 +2278,7 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n \t\tdefault:\n \t\t\t/* Not a delta hence we've already got all we need. */\n \t\t\toe_set_type(entry, entry->in_pack_type);\n-\t\t\tSET_SIZE(entry, in_pack_size);\n+\t\t\tSET_SIZE(entry, cast_size_t_to_ulong(in_pack_size));\n \t\t\tentry->in_pack_header_size = used;\n \t\t\tif (oe_type(entry) < OBJ_COMMIT || oe_type(entry) > OBJ_BLOB)\n \t\t\t\tgoto give_up;\n@@ -2324,8 +2332,8 @@ static void check_object(struct object_entry *entry, uint32_t object_index)\n \t\tif (have_base &&\n \t\t    can_reuse_delta(&base_ref, entry, &base_entry)) {\n \t\t\toe_set_type(entry, entry->in_pack_type);\n-\t\t\tSET_SIZE(entry, in_pack_size); /* delta size */\n-\t\t\tSET_DELTA_SIZE(entry, in_pack_size);\n+\t\t\tSET_SIZE(entry, cast_size_t_to_ulong(in_pack_size)); /* delta size */\n+\t\t\tSET_DELTA_SIZE(entry, cast_size_t_to_ulong(in_pack_size));\n \n \t\t\tif (base_entry) {\n \t\t\t\tSET_DELTA(entry, base_entry);\n@@ -2734,16 +2742,18 @@ unsigned long oe_get_size_slow(struct packing_data *pack,\n \tstruct pack_window *w_curs;\n \tunsigned char *buf;\n \tenum object_type type;\n-\tunsigned long used, avail, size;\n+\tunsigned long used, avail;\n+\tsize_t size;\n \n \tif (e->type_ != OBJ_OFS_DELTA && e->type_ != OBJ_REF_DELTA) {\n+\t\tunsigned long sz;\n \t\tpacking_data_lock(&to_pack);\n \t\tif (odb_read_object_info(the_repository->objects,\n-\t\t\t\t\t &e->idx.oid, &size) < 0)\n+\t\t\t\t\t &e->idx.oid, &sz) < 0)\n \t\t\tdie(_(\"unable to get size of %s\"),\n \t\t\t    oid_to_hex(&e->idx.oid));\n \t\tpacking_data_unlock(&to_pack);\n-\t\treturn size;\n+\t\treturn sz;\n \t}\n \n \tp = oe_in_pack(pack, e);\n@@ -2760,7 +2770,7 @@ unsigned long oe_get_size_slow(struct packing_data *pack,\n \n \tunuse_pack(&w_curs);\n \tpacking_data_unlock(&to_pack);\n-\treturn size;\n+\treturn cast_size_t_to_ulong(size);\n }\n \n static int try_delta(struct unpacked *trg, struct unpacked *src,\ndiff --git a/object-file.c b/object-file.c\nindex 086b2b65ff..0be2981c7a 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2326,6 +2326,7 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct odb_loose_read_stream *st;\n \tunsigned long mapsize;\n+\tunsigned long size_ul;\n \tvoid *mapped;\n \n \tmapped = odb_source_loose_map_object(source, oid, &mapsize);\n@@ -2349,11 +2350,18 @@ int odb_source_loose_read_object_stream(struct odb_read_stream **out,\n \t\tgoto error;\n \t}\n \n-\toi.sizep = &st->base.size;\n+\t/*\n+\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n+\t * st->base.size is size_t (64-bit). Use temporary variable.\n+\t * Note: loose objects >4GB would still truncate here, but such\n+\t * large loose objects are uncommon (they'd normally be packed).\n+\t */\n+\toi.sizep = &size_ul;\n \toi.typep = &st->base.type;\n \n \tif (parse_loose_header(st->hdr, &oi) < 0 || st->base.type < 0)\n \t\tgoto error;\n+\tst->base.size = size_ul;\n \n \tst->mapped = mapped;\n \tst->mapsize = mapsize;\ndiff --git a/odb/streaming.c b/odb/streaming.c\nindex 5927a12954..af2adf5ce7 100644\n--- a/odb/streaming.c\n+++ b/odb/streaming.c\n@@ -157,15 +157,26 @@ static int open_istream_incore(struct odb_read_stream **out,\n \t\t.base.read = read_istream_incore,\n \t};\n \tstruct odb_incore_read_stream *st;\n+\tunsigned long size_ul;\n \tint ret;\n \n \toi.typep = &stream.base.type;\n-\toi.sizep = &stream.base.size;\n+\t/*\n+\t * object_info.sizep is unsigned long* (32-bit on Windows), but\n+\t * stream.base.size is size_t (64-bit). We use a temporary variable\n+\t * because the types are incompatible. Note: this path still truncates\n+\t * for >4GB objects, but large objects should use pack streaming\n+\t * (packfile_store_read_object_stream) which handles size_t properly.\n+\t * This incore fallback is only used for small objects or when pack\n+\t * streaming is unavailable.\n+\t */\n+\toi.sizep = &size_ul;\n \toi.contentp = (void **)&stream.buf;\n \tret = odb_read_object_info_extended(odb, oid, &oi,\n \t\t\t\t\t    OBJECT_INFO_DIE_IF_CORRUPT);\n \tif (ret)\n \t\treturn ret;\n+\tstream.base.size = size_ul;\n \n \tCALLOC_ARRAY(st, 1);\n \t*st = stream;\ndiff --git a/odb/streaming.h b/odb/streaming.h\nindex c7861f7e13..517e2ea2d3 100644\n--- a/odb/streaming.h\n+++ b/odb/streaming.h\n@@ -21,7 +21,7 @@ struct odb_read_stream {\n \todb_read_stream_close_fn close;\n \todb_read_stream_read_fn read;\n \tenum object_type type;\n-\tunsigned long size; /* inflated size of full object */\n+\tsize_t size; /* inflated size of full object */\n };\n \n /*\ndiff --git a/oss-fuzz/fuzz-pack-headers.c b/oss-fuzz/fuzz-pack-headers.c\nindex 150c0f5fa2..ef61ab577c 100644\n--- a/oss-fuzz/fuzz-pack-headers.c\n+++ b/oss-fuzz/fuzz-pack-headers.c\n@@ -6,7 +6,7 @@ int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size);\n int LLVMFuzzerTestOneInput(const uint8_t *data, size_t size)\n {\n \tenum object_type type;\n-\tunsigned long len;\n+\tsize_t len;\n \n \tunpack_object_header_buffer((const unsigned char *)data,\n \t\t\t\t    (unsigned long)size, &type, &len);\ndiff --git a/pack-bitmap.c b/pack-bitmap.c\nindex f6ec18d83a..f9af8a96bd 100644\n--- a/pack-bitmap.c\n+++ b/pack-bitmap.c\n@@ -2270,7 +2270,7 @@ static int try_partial_reuse(struct bitmap_index *bitmap_git,\n {\n \toff_t delta_obj_offset;\n \tenum object_type type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \tif (pack_pos >= pack->p->num_objects)\n \t\treturn -1; /* not actually in the pack */\ndiff --git a/pack-check.c b/pack-check.c\nindex 79992bb509..2792f34d25 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -110,7 +110,7 @@ static int verify_packfile(struct repository *r,\n \t\tvoid *data;\n \t\tstruct object_id oid;\n \t\tenum object_type type;\n-\t\tunsigned long size;\n+\t\tsize_t size;\n \t\toff_t curpos;\n \t\tint data_valid;\n \n@@ -143,7 +143,9 @@ static int verify_packfile(struct repository *r,\n \t\t\tdata = NULL;\n \t\t\tdata_valid = 0;\n \t\t} else {\n-\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &size);\n+\t\t\tunsigned long sz;\n+\t\t\tdata = unpack_entry(r, p, entries[i].offset, &type, &sz);\n+\t\t\tsize = sz;\n \t\t\tdata_valid = 1;\n \t\t}\n \ndiff --git a/packfile.c b/packfile.c\nindex b012d648ad..fdae91dd11 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1133,7 +1133,7 @@ out:\n }\n \n unsigned long unpack_object_header_buffer(const unsigned char *buf,\n-\t\tunsigned long len, enum object_type *type, unsigned long *sizep)\n+\t\tunsigned long len, enum object_type *type, size_t *sizep)\n {\n \tunsigned shift;\n \tsize_t size, c;\n@@ -1144,7 +1144,11 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \tsize = c & 15;\n \tshift = 4;\n \twhile (c & 0x80) {\n-\t\tif (len <= used || (bitsizeof(long) - 7) < shift) {\n+\t\t/*\n+\t\t * Each continuation byte adds 7 bits. Ensure shift won't\n+\t\t * overflow size_t (use size_t not long for 64-bit on Windows).\n+\t\t */\n+\t\tif (len <= used || (bitsizeof(size_t) - 7) < shift) {\n \t\t\terror(\"bad object header\");\n \t\t\tsize = used = 0;\n \t\t\tbreak;\n@@ -1153,7 +1157,7 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \t\tsize = st_add(size, st_left_shift(c & 0x7f, shift));\n \t\tshift += 7;\n \t}\n-\t*sizep = cast_size_t_to_ulong(size);\n+\t*sizep = size;\n \treturn used;\n }\n \n@@ -1215,7 +1219,7 @@ unsigned long get_size_from_delta(struct packed_git *p,\n int unpack_object_header(struct packed_git *p,\n \t\t\t struct pack_window **w_curs,\n \t\t\t off_t *curpos,\n-\t\t\t unsigned long *sizep)\n+\t\t\t size_t *sizep)\n {\n \tunsigned char *base;\n \tunsigned long left;\n@@ -1367,7 +1371,7 @@ static enum object_type packed_to_object_type(struct repository *r,\n \n \twhile (type == OBJ_OFS_DELTA || type == OBJ_REF_DELTA) {\n \t\toff_t base_offset;\n-\t\tunsigned long size;\n+\t\tsize_t size;\n \t\t/* Push the object we're going to leave behind */\n \t\tif (poi_stack_nr >= poi_stack_alloc && poi_stack == small_poi_stack) {\n \t\t\tpoi_stack_alloc = alloc_nr(poi_stack_nr);\n@@ -1586,7 +1590,7 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n \t\t\t\t\t     uint32_t *maybe_index_pos, struct object_info *oi)\n {\n \tstruct pack_window *w_curs = NULL;\n-\tunsigned long size;\n+\tsize_t size;\n \toff_t curpos = obj_offset;\n \tenum object_type type = OBJ_NONE;\n \tuint32_t pack_pos;\n@@ -1778,7 +1782,7 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n \tstruct pack_window *w_curs = NULL;\n \toff_t curpos = obj_offset;\n \tvoid *data = NULL;\n-\tunsigned long size;\n+\tsize_t size;\n \tenum object_type type;\n \tstruct unpack_entry_stack_ent small_delta_stack[UNPACK_ENTRY_STACK_PREALLOC];\n \tstruct unpack_entry_stack_ent *delta_stack = small_delta_stack;\n@@ -1943,8 +1947,10 @@ void *unpack_entry(struct repository *r, struct packed_git *p, off_t obj_offset,\n \t\t\t      (uintmax_t)curpos, p->pack_name);\n \t\t\tdata = NULL;\n \t\t} else {\n+\t\t\tunsigned long sz;\n \t\t\tdata = patch_delta(base, base_size, delta_data,\n-\t\t\t\t\t   delta_size, &size);\n+\t\t\t\t\t   delta_size, &sz);\n+\t\t\tsize = sz;\n \n \t\t\t/*\n \t\t\t * We could not apply the delta; warn the user, but\n@@ -2929,7 +2935,7 @@ int packfile_read_object_stream(struct odb_read_stream **out,\n \tstruct odb_packed_read_stream *stream;\n \tstruct pack_window *window = NULL;\n \tenum object_type in_pack_type;\n-\tunsigned long size;\n+\tsize_t size;\n \n \tin_pack_type = unpack_object_header(pack, &window, &offset, &size);\n \tunuse_pack(&window);\ndiff --git a/packfile.h b/packfile.h\nindex 9b647da7dd..49d6bdecf6 100644\n--- a/packfile.h\n+++ b/packfile.h\n@@ -456,9 +456,9 @@ off_t find_pack_entry_one(const struct object_id *oid, struct packed_git *);\n \n int is_pack_valid(struct packed_git *);\n void *unpack_entry(struct repository *r, struct packed_git *, off_t, enum object_type *, unsigned long *);\n-unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, unsigned long *sizep);\n+unsigned long unpack_object_header_buffer(const unsigned char *buf, unsigned long len, enum object_type *type, size_t *sizep);\n unsigned long get_size_from_delta(struct packed_git *, struct pack_window **, off_t);\n-int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, unsigned long *);\n+int unpack_object_header(struct packed_git *, struct pack_window **, off_t *, size_t *);\n off_t get_delta_base(struct packed_git *p, struct pack_window **w_curs,\n \t\t     off_t *curpos, enum object_type type,\n \t\t     off_t delta_obj_offset);\n-- \ngitgitgadget\n\n"},{"id":"542893","messageId":"8e87a4e71f8684fbd4331b42b3237ffa4284501c.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 04/11] delta, packfile: use size_t for delta header sizes","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:42Z","receivedAt":"2026-05-08T08:16:58Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe delta header decoding functions return unsigned long, which\ntruncates on Windows for objects larger than 4GB. Introduce size_t\nvariants get_delta_hdr_size_sz() and get_size_from_delta_sz() that\npreserve the full 64-bit size, and use them in packed_object_info()\nwhere the size is needed for streaming decisions.\n\nThis was originally authored by LordKiRon <https://github.com/LordKiRon>,\nwho preferred not to reveal their real name and therefore agreed that I\ntake over authorship.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n delta.h    | 14 ++++++++++++--\n packfile.c | 33 ++++++++++++++++++++++++---------\n 2 files changed, 36 insertions(+), 11 deletions(-)\n\ndiff --git a/delta.h b/delta.h\nindex 8a56ec0799..fad68cfc45 100644\n--- a/delta.h\n+++ b/delta.h\n@@ -86,8 +86,11 @@ void *patch_delta(const void *src_buf, unsigned long src_size,\n  * This must be called twice on the delta data buffer, first to get the\n  * expected source buffer size, and again to get the target buffer size.\n  */\n-static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n-\t\t\t\t\t       const unsigned char *top)\n+/*\n+ * Size_t variant that doesn't truncate - use for >4GB objects on Windows.\n+ */\n+static inline size_t get_delta_hdr_size_sz(const unsigned char **datap,\n+\t\t\t\t\t   const unsigned char *top)\n {\n \tconst unsigned char *data = *datap;\n \tsize_t cmd, size = 0;\n@@ -98,6 +101,13 @@ static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n \t\ti += 7;\n \t} while (cmd & 0x80 && data < top);\n \t*datap = data;\n+\treturn size;\n+}\n+\n+static inline unsigned long get_delta_hdr_size(const unsigned char **datap,\n+\t\t\t\t\t       const unsigned char *top)\n+{\n+\tsize_t size = get_delta_hdr_size_sz(datap, top);\n \treturn cast_size_t_to_ulong(size);\n }\n \ndiff --git a/packfile.c b/packfile.c\nindex fdae91dd11..4208f53046 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -1161,9 +1161,12 @@ unsigned long unpack_object_header_buffer(const unsigned char *buf,\n \treturn used;\n }\n \n-unsigned long get_size_from_delta(struct packed_git *p,\n-\t\t\t\t  struct pack_window **w_curs,\n-\t\t\t\t  off_t curpos)\n+/*\n+ * Size_t variant for >4GB delta results on Windows.\n+ */\n+static size_t get_size_from_delta_sz(struct packed_git *p,\n+\t\t\t\t     struct pack_window **w_curs,\n+\t\t\t\t     off_t curpos)\n {\n \tconst unsigned char *data;\n \tunsigned char delta_head[20], *in;\n@@ -1210,10 +1213,18 @@ unsigned long get_size_from_delta(struct packed_git *p,\n \tdata = delta_head;\n \n \t/* ignore base size */\n-\tget_delta_hdr_size(&data, delta_head+sizeof(delta_head));\n+\tget_delta_hdr_size_sz(&data, delta_head+sizeof(delta_head));\n \n \t/* Read the result size */\n-\treturn get_delta_hdr_size(&data, delta_head+sizeof(delta_head));\n+\treturn get_delta_hdr_size_sz(&data, delta_head+sizeof(delta_head));\n+}\n+\n+unsigned long get_size_from_delta(struct packed_git *p,\n+\t\t\t\t  struct pack_window **w_curs,\n+\t\t\t\t  off_t curpos)\n+{\n+\tsize_t size = get_size_from_delta_sz(p, w_curs, curpos);\n+\treturn cast_size_t_to_ulong(size);\n }\n \n int unpack_object_header(struct packed_git *p,\n@@ -1618,14 +1629,18 @@ static int packed_object_info_with_index_pos(struct packed_git *p, off_t obj_off\n \t\t\t\tret = -1;\n \t\t\t\tgoto out;\n \t\t\t}\n-\t\t\t*oi->sizep = get_size_from_delta(p, &w_curs, tmp_pos);\n-\t\t\tif (*oi->sizep == 0) {\n+\t\t\t/*\n+\t\t\t * Use size_t variant to avoid die() on >4GB deltas.\n+\t\t\t * oi->sizep is unsigned long, so truncation may occur,\n+\t\t\t * but streaming code uses its own size_t tracking.\n+\t\t\t */\n+\t\t\tsize = get_size_from_delta_sz(p, &w_curs, tmp_pos);\n+\t\t\tif (size == 0) {\n \t\t\t\tret = -1;\n \t\t\t\tgoto out;\n \t\t\t}\n-\t\t} else {\n-\t\t\t*oi->sizep = size;\n \t\t}\n+\t\t*oi->sizep = (unsigned long)size;\n \t}\n \n \tif (oi->disk_sizep || (oi->mtimep && p->is_cruft)) {\n-- \ngitgitgadget\n\n"},{"id":"542894","messageId":"34fec4a32d6847441a765c4e91a63afc01fae742.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 05/11] test-tool: add a helper to synthesize large packfiles","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:43Z","receivedAt":"2026-05-08T08:17:00Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nTo test Git's behavior with very large pack files, we need a way to\ngenerate such files quickly.\n\nA naive approach using only readily-available Git commands would take\nover 10 hours for a 4GB pack file, which is prohibitive.\n\nSide-stepping Git's machinery and actual zlib compression by writing\nuncompressed content with the appropriate zlib header makes things\nmuch faster. The fastest method using this approach generates many\nsmall, unreachable blob objects and takes about 1.5 minutes for 4GB.\nHowever, this cannot be used because we need to test git clone, which\nrequires a reachable commit history.\n\nGenerating many reachable commits with small, uncompressed blobs takes\nabout 4 minutes for 4GB. But this approach 1) does not reproduce the\nissues we want to fix (which require individual objects larger than\n4GB) and 2) is comparatively slow because of the many SHA-1\ncalculations.\n\nThe approach taken here generates a single large blob (filled with NUL\nbytes), along with the trees and commits needed to make it reachable.\nThis takes about 2.5 minutes for 4.5GB, which is the fastest option\nthat produces a valid, clonable repository with an object large enough\nto trigger the bugs we want to test.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Makefile                   |   1 +\n compat/zlib-compat.h       |   2 +\n t/helper/meson.build       |   1 +\n t/helper/test-synthesize.c | 250 +++++++++++++++++++++++++++++++++++++\n t/helper/test-tool.c       |   1 +\n t/helper/test-tool.h       |   1 +\n 6 files changed, 256 insertions(+)\n create mode 100644 t/helper/test-synthesize.c\n\ndiff --git a/Makefile b/Makefile\nindex cedc234173..85405cb5b8 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -872,6 +872,7 @@ TEST_BUILTINS_OBJS += test-submodule-config.o\n TEST_BUILTINS_OBJS += test-submodule-nested-repo-config.o\n TEST_BUILTINS_OBJS += test-submodule.o\n TEST_BUILTINS_OBJS += test-subprocess.o\n+TEST_BUILTINS_OBJS += test-synthesize.o\n TEST_BUILTINS_OBJS += test-trace2.o\n TEST_BUILTINS_OBJS += test-truncate.o\n TEST_BUILTINS_OBJS += test-userdiff.o\ndiff --git a/compat/zlib-compat.h b/compat/zlib-compat.h\nindex ac08276622..5078c5ef6c 100644\n--- a/compat/zlib-compat.h\n+++ b/compat/zlib-compat.h\n@@ -7,6 +7,8 @@\n # define z_stream_s zng_stream_s\n # define gz_header_s zng_gz_header_s\n \n+# define adler32(adler, buf, len) zng_adler32(adler, buf, len)\n+\n # define crc32(crc, buf, len) zng_crc32(crc, buf, len)\n \n # define inflate(strm, bits) zng_inflate(strm, bits)\ndiff --git a/t/helper/meson.build b/t/helper/meson.build\nindex 675e64c010..3235f10ab8 100644\n--- a/t/helper/meson.build\n+++ b/t/helper/meson.build\n@@ -69,6 +69,7 @@ test_tool_sources = [\n   'test-submodule-nested-repo-config.c',\n   'test-submodule.c',\n   'test-subprocess.c',\n+  'test-synthesize.c',\n   'test-tool.c',\n   'test-trace2.c',\n   'test-truncate.c',\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nnew file mode 100644\nindex 0000000000..3ce7078078\n--- /dev/null\n+++ b/t/helper/test-synthesize.c\n@@ -0,0 +1,250 @@\n+#define USE_THE_REPOSITORY_VARIABLE\n+\n+#include \"test-tool.h\"\n+#include \"git-compat-util.h\"\n+#include \"git-zlib.h\"\n+#include \"hash.h\"\n+#include \"hex.h\"\n+#include \"object-file.h\"\n+#include \"object.h\"\n+#include \"pack.h\"\n+#include \"parse-options.h\"\n+#include \"parse.h\"\n+#include \"repository.h\"\n+#include \"setup.h\"\n+#include \"strbuf.h\"\n+#include \"write-or-die.h\"\n+\n+#define BLOCK_SIZE 0xffff\n+static const unsigned char zeros[BLOCK_SIZE];\n+\n+/*\n+ * Write data as an uncompressed zlib stream.\n+ * For data larger than 64KB, writes multiple uncompressed blocks.\n+ * If data is NULL, writes zeros.\n+ * Updates the pack checksum context.\n+ */\n+static void write_uncompressed_zlib(FILE *f, struct git_hash_ctx *pack_ctx,\n+\t\t\t\t    const void *data, size_t len,\n+\t\t\t\t    const struct git_hash_algo *algo)\n+{\n+\tunsigned char zlib_header[2] = { 0x78, 0x01 }; /* CMF, FLG */\n+\tunsigned char block_header[5];\n+\tconst unsigned char *p = data;\n+\tsize_t remaining = len;\n+\tuint32_t adler = 1L; /* adler32 initial value */\n+\tunsigned char adler_buf[4];\n+\n+\t/* Write zlib header */\n+\tfwrite_or_die(f, zlib_header, sizeof(zlib_header));\n+\talgo->update_fn(pack_ctx, zlib_header, 2);\n+\n+\t/* Write uncompressed blocks (max 64KB each) */\n+\tdo {\n+\t\tsize_t block_len = remaining > BLOCK_SIZE ? BLOCK_SIZE : remaining;\n+\t\tint is_final = (block_len == remaining);\n+\t\tconst unsigned char *block_data = data ? p : zeros;\n+\n+\t\tblock_header[0] = is_final ? 0x01 : 0x00;\n+\t\tblock_header[1] = block_len & 0xff;\n+\t\tblock_header[2] = (block_len >> 8) & 0xff;\n+\t\tblock_header[3] = block_header[1] ^ 0xff;\n+\t\tblock_header[4] = block_header[2] ^ 0xff;\n+\n+\t\tfwrite_or_die(f, block_header, sizeof(block_header));\n+\t\talgo->update_fn(pack_ctx, block_header, 5);\n+\n+\t\tif (block_len) {\n+\t\t\tfwrite_or_die(f, block_data, block_len);\n+\t\t\talgo->update_fn(pack_ctx, block_data, block_len);\n+\t\t\tadler = adler32(adler, block_data, block_len);\n+\t\t}\n+\n+\t\tif (data)\n+\t\t\tp += block_len;\n+\t\tremaining -= block_len;\n+\t} while (remaining > 0);\n+\n+\t/* Write adler32 checksum */\n+\tput_be32(adler_buf, adler);\n+\tfwrite_or_die(f, adler_buf, sizeof(adler_buf));\n+\talgo->update_fn(pack_ctx, adler_buf, 4);\n+}\n+\n+/*\n+ * Write an uncompressed object to the pack file.\n+ * If `data == NULL`, it is treated like a buffer to NUL bytes.\n+ * Updates the pack checksum context.\n+ */\n+static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n+\t\t\t      enum object_type type,\n+\t\t\t      const void *data, size_t len,\n+\t\t\t      struct object_id *oid,\n+\t\t\t      const struct git_hash_algo *algo)\n+{\n+\tunsigned char pack_header[MAX_PACK_OBJECT_HEADER];\n+\tchar object_header[32];\n+\tint pack_header_len, object_header_len;\n+\tstruct git_hash_ctx ctx;\n+\n+\t/* Write pack object header */\n+\tpack_header_len = encode_in_pack_object_header(pack_header,\n+\t\t\t\t\t\t       sizeof(pack_header),\n+\t\t\t\t\t\t       type, len);\n+\tfwrite_or_die(f, pack_header, pack_header_len);\n+\talgo->update_fn(pack_ctx, pack_header, pack_header_len);\n+\n+\t/* Write the data as uncompressed zlib */\n+\twrite_uncompressed_zlib(f, pack_ctx, data, len, algo);\n+\n+\talgo->init_fn(&ctx);\n+\tobject_header_len = format_object_header(object_header,\n+\t\t\t\t\t\t sizeof(object_header),\n+\t\t\t\t\t\t type, len);\n+\talgo->update_fn(&ctx, object_header, object_header_len);\n+\tif (data)\n+\t\talgo->update_fn(&ctx, data, len);\n+\telse {\n+\t\tfor (size_t i = len / BLOCK_SIZE; i; i--)\n+\t\t\talgo->update_fn(&ctx, zeros, BLOCK_SIZE);\n+\t\talgo->update_fn(&ctx, zeros, len % BLOCK_SIZE);\n+\t}\n+\talgo->final_oid_fn(oid, &ctx);\n+}\n+\n+/*\n+ * Generate a pack file with a single large (>4GB) reachable object.\n+ *\n+ * Creates:\n+ *   1. A large blob (all NUL bytes)\n+ *   2. A tree containing that blob as \"file\"\n+ *   3. A commit using that tree\n+ *   4. The empty tree\n+ *   5. A child commit using the empty tree\n+ *\n+ * This is useful for testing that Git can handle objects larger than 4GB.\n+ */\n+static int generate_pack_with_large_object(const char *path, size_t blob_size,\n+\t\t\t\t\t   const struct git_hash_algo *algo)\n+{\n+\tFILE *f = xfopen(path, \"wb\");\n+\tstruct git_hash_ctx pack_ctx;\n+\tunsigned char pack_hash[GIT_MAX_RAWSZ];\n+\tstruct object_id blob_oid, tree_oid, commit_oid, empty_tree_oid, final_commit_oid;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tconst uint32_t object_count = 5;\n+\tstruct pack_header pack_header = {\n+\t\t.hdr_signature = htonl(PACK_SIGNATURE),\n+\t\t.hdr_version = htonl(PACK_VERSION),\n+\t\t.hdr_entries = htonl(object_count),\n+\t};\n+\n+\talgo->init_fn(&pack_ctx);\n+\n+\t/* Write pack header */\n+\tfwrite_or_die(f, &pack_header, sizeof(pack_header));\n+\talgo->update_fn(&pack_ctx, &pack_header, sizeof(pack_header));\n+\n+\t/* 1. Write the large blob */\n+\twrite_pack_object(f, &pack_ctx, OBJ_BLOB, NULL, blob_size, &blob_oid, algo);\n+\n+\t/* 2. Write tree containing the blob as \"file\" */\n+\tstrbuf_addf(&buf, \"100644 file%c\", '\\0');\n+\tstrbuf_add(&buf, blob_oid.hash, algo->rawsz);\n+\twrite_pack_object(f, &pack_ctx, OBJ_TREE, buf.buf, buf.len, &tree_oid, algo);\n+\n+\t/* 3. Write commit using that tree */\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf,\n+\t\t    \"tree %s\\n\"\n+\t\t    \"author A U Thor <author@example.com> 1234567890 +0000\\n\"\n+\t\t    \"committer C O Mitter <committer@example.com> 1234567890 +0000\\n\"\n+\t\t    \"\\n\"\n+\t\t    \"Large blob commit\\n\",\n+\t\t    oid_to_hex(&tree_oid));\n+\twrite_pack_object(f, &pack_ctx, OBJ_COMMIT, buf.buf, buf.len, &commit_oid, algo);\n+\n+\t/* 4. Write the empty tree */\n+\twrite_pack_object(f, &pack_ctx, OBJ_TREE, \"\", 0, &empty_tree_oid, algo);\n+\n+\t/* 5. Write final commit using empty tree, with previous commit as parent */\n+\tstrbuf_reset(&buf);\n+\tstrbuf_addf(&buf,\n+\t\t    \"tree %s\\n\"\n+\t\t    \"parent %s\\n\"\n+\t\t    \"author A U Thor <author@example.com> 1234567890 +0000\\n\"\n+\t\t    \"committer C O Mitter <committer@example.com> 1234567890 +0000\\n\"\n+\t\t    \"\\n\"\n+\t\t    \"Empty tree commit\\n\",\n+\t\t    oid_to_hex(&empty_tree_oid),\n+\t\t    oid_to_hex(&commit_oid));\n+\twrite_pack_object(f, &pack_ctx, OBJ_COMMIT, buf.buf, buf.len, &final_commit_oid, algo);\n+\n+\t/* Write pack trailer (checksum) */\n+\talgo->final_fn(pack_hash, &pack_ctx);\n+\tfwrite_or_die(f, pack_hash, algo->rawsz);\n+\tif (fclose(f))\n+\t\tdie_errno(_(\"could not close '%s'\"), path);\n+\n+\tstrbuf_release(&buf);\n+\n+\t/* Print the final commit OID so caller can set up refs */\n+\tprintf(\"%s\\n\", oid_to_hex(&final_commit_oid));\n+\n+\treturn 0;\n+}\n+\n+static int cmd__synthesize__pack(int argc, const char **argv,\n+\t\t\t\t const char *prefix UNUSED,\n+\t\t\t\t struct repository *repo)\n+{\n+\tint non_git;\n+\tint reachable_large = 0;\n+\tconst struct git_hash_algo *algo;\n+\tsize_t blob_size;\n+\tuintmax_t blob_size_u;\n+\tconst char *path;\n+\tconst char * const usage[] = {\n+\t\t\"test-tool synthesize pack \"\n+\t\t\"--reachable-large <blob-size> <filename>\",\n+\t\tNULL\n+\t};\n+\tstruct option options[] = {\n+\t\tOPT_BOOL(0, \"reachable-large\", &reachable_large,\n+\t\t\t N_(\"write a pack with a single reachable large blob\")),\n+\t\tOPT_END()\n+\t};\n+\n+\tsetup_git_directory_gently(&non_git);\n+\trepo = the_repository;\n+\talgo = repo->hash_algo;\n+\n+\targc = parse_options(argc, argv, NULL, options, usage,\n+\t\t\t     PARSE_OPT_KEEP_ARGV0);\n+\tif (argc != 3 || !reachable_large)\n+\t\tusage_with_options(usage, options);\n+\n+\tif (!git_parse_unsigned(argv[1], &blob_size_u,\n+\t\t\t\tmaximum_unsigned_value_of_type(size_t)))\n+\t\tdie(_(\"'%s' is not a valid blob size\"), argv[1]);\n+\tblob_size = blob_size_u;\n+\tpath = argv[2];\n+\n+\treturn !!generate_pack_with_large_object(path, blob_size, algo);\n+}\n+\n+int cmd__synthesize(int argc, const char **argv)\n+{\n+\tconst char *prefix = NULL;\n+\tchar const * const synthesize_usage[] = {\n+\t\t\"test-tool synthesize pack <options>\",\n+\t\tNULL,\n+\t};\n+\tparse_opt_subcommand_fn *fn = NULL;\n+\tstruct option options[] = {\n+\t\tOPT_SUBCOMMAND(\"pack\", &fn, cmd__synthesize__pack),\n+\t\tOPT_END()\n+\t};\n+\targc = parse_options(argc, argv, prefix, options, synthesize_usage, 0);\n+\treturn !!fn(argc, argv, prefix, NULL);\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex a7abc618b3..b71a22b43b 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -82,6 +82,7 @@ static struct test_cmd cmds[] = {\n \t{ \"submodule-config\", cmd__submodule_config },\n \t{ \"submodule-nested-repo-config\", cmd__submodule_nested_repo_config },\n \t{ \"subprocess\", cmd__subprocess },\n+\t{ \"synthesize\", cmd__synthesize },\n \t{ \"trace2\", cmd__trace2 },\n \t{ \"truncate\", cmd__truncate },\n \t{ \"userdiff\", cmd__userdiff },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex 7f150fa1eb..f2885b33d5 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -75,6 +75,7 @@ int cmd__submodule(int argc, const char **argv);\n int cmd__submodule_config(int argc, const char **argv);\n int cmd__submodule_nested_repo_config(int argc, const char **argv);\n int cmd__subprocess(int argc, const char **argv);\n+int cmd__synthesize(int argc, const char **argv);\n int cmd__trace2(int argc, const char **argv);\n int cmd__truncate(int argc, const char **argv);\n int cmd__userdiff(int argc, const char **argv);\n-- \ngitgitgadget\n\n"},{"id":"542895","messageId":"88f992903f8c0f01335f7573e2806ecfba4b508f.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 06/11] t5608: add regression test for >4GB object clone","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:44Z","receivedAt":"2026-05-08T08:17:01Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe shift overflow bug in index-pack and unpack-objects caused incorrect\nobject size calculation when the encoded size required more than 32 bits\nof shift. This would result in corrupted or failed unpacking of objects\nlarger than 4GB.\n\nAdd a test that creates a pack file containing a 4GB+ blob using the\nnew 'test-tool synthesize pack --reachable-large' command, then clones\nthe repository to verify the fix works correctly.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t5608-clone-2gb.sh | 37 +++++++++++++++++++++++++++++++++++++\n 1 file changed, 37 insertions(+)\n\ndiff --git a/t/t5608-clone-2gb.sh b/t/t5608-clone-2gb.sh\nindex 87a8cd9f98..af93302dde 100755\n--- a/t/t5608-clone-2gb.sh\n+++ b/t/t5608-clone-2gb.sh\n@@ -49,4 +49,41 @@ test_expect_success 'clone - with worktree, file:// protocol' '\n \n '\n \n+test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n+\tlarge_blob_size=$((4*1024*1024*1024+1)) &&\n+\tgit init --bare 4gb-repo &&\n+\thead_oid=$(test-tool synthesize pack \\\n+\t\t--reachable-large \"$large_blob_size\" \\\n+\t\t4gb-repo/objects/pack/test.pack) &&\n+\tgit -C 4gb-repo index-pack objects/pack/test.pack &&\n+\tgit -C 4gb-repo update-ref refs/heads/main $head_oid &&\n+\tgit -C 4gb-repo symbolic-ref HEAD refs/heads/main\n+'\n+\n+test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n+\t# The synthesized pack has five objects, so a large unpack limit keeps\n+\t# fetch-pack on the unpack-objects path.\n+\tgit -c fetch.unpackLimit=100 clone --bare \\\n+\t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-unpack &&\n+\n+\t# Verify the large blob survived the clone by comparing its OID\n+\t# between source and clone.  We cannot use \"cat-file -s\" because\n+\t# object_info.sizep is still unsigned long, which truncates >4GB\n+\t# sizes on Windows.  OID equality proves content integrity since\n+\t# the clone already verified checksums via index-pack/unpack-objects.\n+\tsource_blob=$(git -C 4gb-repo rev-parse main^:file) &&\n+\tclone_blob=$(git -C 4gb-clone-unpack rev-parse main^:file) &&\n+\ttest \"$source_blob\" = \"$clone_blob\"\n+'\n+\n+test_expect_success SIZE_T_IS_64BIT 'clone with >4GB object via index-pack' '\n+\t# Force fetch-pack to hand the pack to index-pack instead.\n+\tgit -c fetch.unpackLimit=1 clone --bare \\\n+\t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-index &&\n+\n+\tsource_blob=$(git -C 4gb-repo rev-parse main^:file) &&\n+\tclone_blob=$(git -C 4gb-clone-index rev-parse main^:file) &&\n+\ttest \"$source_blob\" = \"$clone_blob\"\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"542896","messageId":"4f207c8a470af1f8cf00c704043dfa94e6e1420d.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 07/11] test-tool synthesize: use the unsafe hash for speed","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:45Z","receivedAt":"2026-05-08T08:17:04Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nJeff King pointed out on the mailing list [1] that t5608's new >4GB\ntest cases dominate the entire test suite runtime: 160 seconds on his\nlaptop when the rest of the suite finishes in under 90 seconds, and\n305-850 seconds across CI jobs. The bottleneck is that the synthesize\nhelper hashes roughly 8 GB of data through SHA-1 (4 GB for the pack\nchecksum plus 4 GB for the blob OID) for a 4 GB+1 blob.\n\nSince the helper generates known test data, collision detection is\nunnecessary. Switch from repo->hash_algo to unsafe_hash_algo(), which\nuses hardware-accelerated SHA-1 (via OpenSSL or Apple CommonCrypto)\nwhen available.\n\nBenchmarks on an x86_64 machine generating a 4 GB+1 pack (2 runs\neach, interleaved):\n\n  SHA-1 backend      Run 1    Run 2\n  SHA1DC (safe)       75s      80s\n  OpenSSL (unsafe)    21s      19s\n\nThe effect scales linearly. At 64 MB with 10 randomized interleaved\nruns, the OpenSSL unsafe backend shows a 5.4x improvement (median\n0.202s vs 1.088s) with tight variance (stdev 0.028s vs 0.095s).\n\nThe speedup is only realized when the build has a fast unsafe backend\ncompiled in. The CI's linux-TEST-vars job already sets\nOPENSSL_SHA1_UNSAFE=YesPlease; macOS benefits from Apple CommonCrypto\nwhen configured. On builds without a separate unsafe backend (such as\nthe default Windows builds), unsafe_hash_algo() returns the regular\ncollision-detecting implementation and the change is a no-op.\n\n[1] https://lore.kernel.org/git/20260501063805.GA2038915@coredump.intra.peff.net/\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/helper/test-synthesize.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nindex 3ce7078078..e2faaad7b4 100644\n--- a/t/helper/test-synthesize.c\n+++ b/t/helper/test-synthesize.c\n@@ -217,7 +217,7 @@ static int cmd__synthesize__pack(int argc, const char **argv,\n \n \tsetup_git_directory_gently(&non_git);\n \trepo = the_repository;\n-\talgo = repo->hash_algo;\n+\talgo = unsafe_hash_algo(repo->hash_algo);\n \n \targc = parse_options(argc, argv, NULL, options, usage,\n \t\t\t     PARSE_OPT_KEEP_ARGV0);\n-- \ngitgitgadget\n\n"},{"id":"542897","messageId":"2751c21c6e730ac7f0ce9c2606bbd8b56dede11e.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 08/11] test-tool synthesize: precompute pack for 4 GiB + 1","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:46Z","receivedAt":"2026-05-08T08:17:06Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nThe synthesize helper hashes roughly 8 GiB of data through SHA-1 to\nproduce a 4 GiB + 1 pack (4 GiB for the pack checksum, 4 GiB for\nthe blob OID). Since the blob content is all NUL bytes, every byte\nin the resulting pack file is deterministic for a given blob size and\nhash algorithm.\n\nAdd a fast path that writes the pack from precomputed constants:\na 25-byte prefix (pack header, object header, zlib header, first\nblock header), the zero-filled bulk with periodic 5-byte deflate\nblock headers, and a 513-byte suffix (tree, two commits, empty tree,\npack SHA-1 checksum). This eliminates all SHA-1 and adler32\ncomputation, making the helper purely I/O-bound.\n\nThe precomputed constants are stored in a struct fast_pack array\nkeyed by hash algorithm format_id, so that adding SHA-256 support\nlater requires only adding another array entry with its suffix.\n\nThe constants were generated by running the generic path and\nextracting the non-zero bytes from the resulting pack file.\n\nBenchmarks generating a 4 GiB + 1 pack (3 runs each, SHA1DC on\nx86_64):\n\n  generic path:   88s / 81s / 140s\n  fast path:      14s / 13s / 15s\n\nOn CI, where t5608 currently takes 200-850 seconds depending on the\njob, the fast path cuts the pack-generation phase from minutes to\nseconds, leaving only the clone operations themselves.\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/helper/test-synthesize.c | 202 ++++++++++++++++++++++++++++++++++++-\n 1 file changed, 201 insertions(+), 1 deletion(-)\n\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nindex e2faaad7b4..83c40ee02a 100644\n--- a/t/helper/test-synthesize.c\n+++ b/t/helper/test-synthesize.c\n@@ -112,6 +112,201 @@ static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n \talgo->final_oid_fn(oid, &ctx);\n }\n \n+/*\n+ * Fast path: precomputed pack data for a 4 GiB + 1 all-NUL blob.\n+ *\n+ * The generated pack is almost entirely zeros with a small constant\n+ * prefix, periodic deflate block headers, and a constant suffix\n+ * containing the tree, two commits, and the pack checksum.  Because\n+ * every byte is deterministic for a given blob size and hash algorithm,\n+ * we can write the pack without computing any hashes at all, reducing\n+ * runtime from minutes of hash computation to seconds of pure I/O.\n+ *\n+ * The blob is stored as an uncompressed deflate stream: a two-byte\n+ * zlib header, then 65538 blocks of up to 0xffff bytes each, followed\n+ * by an adler32 checksum.  The pack header and deflate framing are\n+ * shared across hash algorithms; only the suffix (which contains OIDs\n+ * and the pack checksum) differs.\n+ *\n+ * Constants were generated by running the generic path and extracting\n+ * the non-zero bytes from the resulting pack file.\n+ */\n+\n+#define FAST_PACK_4G1_BLOB_SIZE ((size_t)4 * 1024 * 1024 * 1024 + 1)\n+#define FAST_PACK_4G1_N_FULL_BLOCKS 65537\n+\n+/*\n+ * Per-hash-algorithm constants for the fast path.  The prefix and\n+ * deflate block structure are identical across algorithms; only the\n+ * suffix (tree, commits, pack checksum) and the commit OID differ.\n+ */\n+struct fast_pack {\n+\tuint32_t format_id;\n+\tconst unsigned char *suffix;\n+\tsize_t suffix_len;\n+\tconst char *commit_oid;\n+};\n+\n+/* Pack header + pack object header + zlib header + first block header */\n+static const unsigned char fast_pack_prefix[] = {\n+\t/* PACK header: signature, version 2, 5 objects */\n+\t0x50, 0x41, 0x43, 0x4b, 0x00, 0x00, 0x00, 0x02,\n+\t0x00, 0x00, 0x00, 0x05,\n+\t/* pack object header: blob, size = 4294967297 */\n+\t0xb1, 0x80, 0x80, 0x80, 0x80, 0x01,\n+\t/* zlib header: CMF=0x78, FLG=0x01 */\n+\t0x78, 0x01,\n+\t/* first non-final block header: BFINAL=0, LEN=0xffff, NLEN=0x0000 */\n+\t0x00, 0xff, 0xff, 0x00, 0x00\n+};\n+\n+/* Every non-final deflate block header is identical */\n+static const unsigned char fast_pack_block_header[] = {\n+\t0x00, 0xff, 0xff, 0x00, 0x00\n+};\n+\n+/* Final block (2 data bytes) + adler32 of 4294967297 NUL bytes */\n+static const unsigned char fast_pack_final_block[] = {\n+\t/* BFINAL=1, LEN=2, NLEN=0xfffd */\n+\t0x01, 0x02, 0x00, 0xfd, 0xff,\n+\t/* 2 NUL data bytes */\n+\t0x00, 0x00,\n+\t/* adler32 */\n+\t0x00, 0xe2, 0x00, 0x01\n+};\n+\n+/*\n+ * SHA-1 suffix: tree, commit, empty tree, final commit, pack checksum.\n+ */\n+static const unsigned char fast_pack_sha1_suffix[] = {\n+\t0xa0, 0x02, 0x78, 0x01, 0x01, 0x20, 0x00, 0xdf,\n+\t0xff, 0x31, 0x30, 0x30, 0x36, 0x34, 0x34, 0x20,\n+\t0x66, 0x69, 0x6c, 0x65, 0x00, 0x3e, 0xb7, 0xfe,\n+\t0xb1, 0x41, 0x3c, 0x75, 0x7f, 0x0d, 0x81, 0x81,\n+\t0xde, 0xb2, 0x8d, 0x1d, 0xab, 0x03, 0xd6, 0x48,\n+\t0x46, 0xb4, 0xb4, 0x0c, 0x60, 0x95, 0x0b, 0x78,\n+\t0x01, 0x01, 0xb5, 0x00, 0x4a, 0xff, 0x74, 0x72,\n+\t0x65, 0x65, 0x20, 0x63, 0x36, 0x38, 0x33, 0x66,\n+\t0x63, 0x63, 0x37, 0x64, 0x31, 0x64, 0x38, 0x33,\n+\t0x65, 0x66, 0x32, 0x66, 0x65, 0x31, 0x61, 0x66,\n+\t0x35, 0x35, 0x32, 0x31, 0x35, 0x64, 0x30, 0x31,\n+\t0x36, 0x38, 0x64, 0x62, 0x35, 0x32, 0x61, 0x33,\n+\t0x61, 0x33, 0x62, 0x0a, 0x61, 0x75, 0x74, 0x68,\n+\t0x6f, 0x72, 0x20, 0x41, 0x20, 0x55, 0x20, 0x54,\n+\t0x68, 0x6f, 0x72, 0x20, 0x3c, 0x61, 0x75, 0x74,\n+\t0x68, 0x6f, 0x72, 0x40, 0x65, 0x78, 0x61, 0x6d,\n+\t0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e,\n+\t0x20, 0x31, 0x32, 0x33, 0x34, 0x35, 0x36, 0x37,\n+\t0x38, 0x39, 0x30, 0x20, 0x2b, 0x30, 0x30, 0x30,\n+\t0x30, 0x0a, 0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74,\n+\t0x74, 0x65, 0x72, 0x20, 0x43, 0x20, 0x4f, 0x20,\n+\t0x4d, 0x69, 0x74, 0x74, 0x65, 0x72, 0x20, 0x3c,\n+\t0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65,\n+\t0x72, 0x40, 0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c,\n+\t0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31,\n+\t0x32, 0x33, 0x34, 0x35, 0x36, 0x37, 0x38, 0x39,\n+\t0x30, 0x20, 0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a,\n+\t0x0a, 0x4c, 0x61, 0x72, 0x67, 0x65, 0x20, 0x62,\n+\t0x6c, 0x6f, 0x62, 0x20, 0x63, 0x6f, 0x6d, 0x6d,\n+\t0x69, 0x74, 0x0a, 0xc6, 0x55, 0x37, 0x6b, 0x20,\n+\t0x78, 0x01, 0x01, 0x00, 0x00, 0xff, 0xff, 0x00,\n+\t0x00, 0x00, 0x01, 0x95, 0x0e, 0x78, 0x01, 0x01,\n+\t0xe5, 0x00, 0x1a, 0xff, 0x74, 0x72, 0x65, 0x65,\n+\t0x20, 0x34, 0x62, 0x38, 0x32, 0x35, 0x64, 0x63,\n+\t0x36, 0x34, 0x32, 0x63, 0x62, 0x36, 0x65, 0x62,\n+\t0x39, 0x61, 0x30, 0x36, 0x30, 0x65, 0x35, 0x34,\n+\t0x62, 0x66, 0x38, 0x64, 0x36, 0x39, 0x32, 0x38,\n+\t0x38, 0x66, 0x62, 0x65, 0x65, 0x34, 0x39, 0x30,\n+\t0x34, 0x0a, 0x70, 0x61, 0x72, 0x65, 0x6e, 0x74,\n+\t0x20, 0x63, 0x35, 0x62, 0x32, 0x31, 0x63, 0x36,\n+\t0x31, 0x31, 0x61, 0x61, 0x35, 0x39, 0x34, 0x65,\n+\t0x63, 0x39, 0x66, 0x64, 0x37, 0x65, 0x39, 0x32,\n+\t0x63, 0x66, 0x39, 0x36, 0x34, 0x38, 0x39, 0x31,\n+\t0x34, 0x63, 0x61, 0x34, 0x63, 0x32, 0x34, 0x31,\n+\t0x32, 0x0a, 0x61, 0x75, 0x74, 0x68, 0x6f, 0x72,\n+\t0x20, 0x41, 0x20, 0x55, 0x20, 0x54, 0x68, 0x6f,\n+\t0x72, 0x20, 0x3c, 0x61, 0x75, 0x74, 0x68, 0x6f,\n+\t0x72, 0x40, 0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c,\n+\t0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31,\n+\t0x32, 0x33, 0x34, 0x35, 0x36, 0x37, 0x38, 0x39,\n+\t0x30, 0x20, 0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a,\n+\t0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65,\n+\t0x72, 0x20, 0x43, 0x20, 0x4f, 0x20, 0x4d, 0x69,\n+\t0x74, 0x74, 0x65, 0x72, 0x20, 0x3c, 0x63, 0x6f,\n+\t0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65, 0x72, 0x40,\n+\t0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e,\n+\t0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33,\n+\t0x34, 0x35, 0x36, 0x37, 0x38, 0x39, 0x30, 0x20,\n+\t0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a, 0x0a, 0x45,\n+\t0x6d, 0x70, 0x74, 0x79, 0x20, 0x74, 0x72, 0x65,\n+\t0x65, 0x20, 0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74,\n+\t0x0a, 0xaa, 0xb8, 0x45, 0x01, 0x8e, 0xfc, 0xf0,\n+\t0x2f, 0x9c, 0xc5, 0xcc, 0x4f, 0x6a, 0x1a, 0xc9,\n+\t0x2b, 0x23, 0xa9, 0xff, 0x91, 0x06, 0xc2, 0x70,\n+\t0xe3\n+};\n+\n+static const struct fast_pack fast_packs[] = {\n+\t{\n+\t\t.format_id = GIT_SHA1_FORMAT_ID,\n+\t\t.suffix = fast_pack_sha1_suffix,\n+\t\t.suffix_len = sizeof(fast_pack_sha1_suffix),\n+\t\t.commit_oid = \"aac43daf40d0377af31aa9c798a4ae8a31b55c1d\",\n+\t},\n+};\n+\n+/*\n+ * Try the fast path for known blob sizes.  Returns 1 if the pack was\n+ * written from precomputed constants, 0 if the caller should fall\n+ * through to the generic path.\n+ */\n+static int generate_fast_pack(const char *path, size_t blob_size,\n+\t\t\t      const struct git_hash_algo *algo)\n+{\n+\tconst struct fast_pack *fp = NULL;\n+\tFILE *f;\n+\tsize_t i;\n+\n+\tif (blob_size != FAST_PACK_4G1_BLOB_SIZE)\n+\t\treturn 0;\n+\n+\tfor (i = 0; i < ARRAY_SIZE(fast_packs); i++) {\n+\t\tif (fast_packs[i].format_id == algo->format_id) {\n+\t\t\tfp = &fast_packs[i];\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\tif (!fp)\n+\t\treturn 0;\n+\n+\tf = xfopen(path, \"wb\");\n+\n+\tfwrite_or_die(f, fast_pack_prefix, sizeof(fast_pack_prefix));\n+\n+\t/* First full block: 0xffff zero bytes (header already in prefix) */\n+\tfwrite_or_die(f, zeros, BLOCK_SIZE);\n+\n+\t/* Remaining non-final full blocks */\n+\tfor (i = 1; i < FAST_PACK_4G1_N_FULL_BLOCKS; i++) {\n+\t\tfwrite_or_die(f, fast_pack_block_header,\n+\t\t\t      sizeof(fast_pack_block_header));\n+\t\tfwrite_or_die(f, zeros, BLOCK_SIZE);\n+\t}\n+\n+\t/* Final block (2 data bytes) + adler32 */\n+\tfwrite_or_die(f, fast_pack_final_block,\n+\t\t      sizeof(fast_pack_final_block));\n+\n+\t/* Tree, commits, and pack checksum */\n+\tfwrite_or_die(f, fp->suffix, fp->suffix_len);\n+\n+\tif (fclose(f))\n+\t\tdie_errno(_(\"could not close '%s'\"), path);\n+\n+\tprintf(\"%s\\n\", fp->commit_oid);\n+\treturn 1;\n+}\n+\n /*\n  * Generate a pack file with a single large (>4GB) reachable object.\n  *\n@@ -127,7 +322,7 @@ static void write_pack_object(FILE *f, struct git_hash_ctx *pack_ctx,\n static int generate_pack_with_large_object(const char *path, size_t blob_size,\n \t\t\t\t\t   const struct git_hash_algo *algo)\n {\n-\tFILE *f = xfopen(path, \"wb\");\n+\tFILE *f;\n \tstruct git_hash_ctx pack_ctx;\n \tunsigned char pack_hash[GIT_MAX_RAWSZ];\n \tstruct object_id blob_oid, tree_oid, commit_oid, empty_tree_oid, final_commit_oid;\n@@ -139,6 +334,11 @@ static int generate_pack_with_large_object(const char *path, size_t blob_size,\n \t\t.hdr_entries = htonl(object_count),\n \t};\n \n+\tif (generate_fast_pack(path, blob_size, algo))\n+\t\treturn 0;\n+\n+\tf = xfopen(path, \"wb\");\n+\n \talgo->init_fn(&pack_ctx);\n \n \t/* Write pack header */\n-- \ngitgitgadget\n\n"},{"id":"542898","messageId":"3a006d96c3e7c41c5ac4a621d1d4db2d7d1c91ff.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 09/11] test-tool synthesize: add precomputed SHA-256 pack for 4 GiB + 1","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:47Z","receivedAt":"2026-05-08T08:17:07Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nAdd a SHA-256 entry to the fast_packs[] table. The pack prefix and\ndeflate block structure are identical to SHA-1 (the pack format does\nnot encode the hash algorithm in its header). Only the suffix differs:\nSHA-256 OIDs are 32 bytes instead of 20, giving a 609-byte suffix\ncompared to 513 for SHA-1, and a different pack checksum.\n\nThe constants were generated by running the generic path inside a\nrepository initialized with --object-format=sha256.\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/helper/test-synthesize.c | 91 ++++++++++++++++++++++++++++++++++++++\n 1 file changed, 91 insertions(+)\n\ndiff --git a/t/helper/test-synthesize.c b/t/helper/test-synthesize.c\nindex 83c40ee02a..1f28ecf0f2 100644\n--- a/t/helper/test-synthesize.c\n+++ b/t/helper/test-synthesize.c\n@@ -246,6 +246,90 @@ static const unsigned char fast_pack_sha1_suffix[] = {\n \t0xe3\n };\n \n+/*\n+ * SHA-256 suffix: same structure, but with 32-byte OIDs and SHA-256\n+ * pack checksum (609 bytes vs 513 for SHA-1).\n+ */\n+static const unsigned char fast_pack_sha256_suffix[] = {\n+\t0xac, 0x02, 0x78, 0x01, 0x01, 0x2c, 0x00, 0xd3,\n+\t0xff, 0x31, 0x30, 0x30, 0x36, 0x34, 0x34, 0x20,\n+\t0x66, 0x69, 0x6c, 0x65, 0x00, 0x42, 0x53, 0xc1,\n+\t0x8a, 0x9f, 0x5e, 0xc3, 0xbb, 0x47, 0xb0, 0x83,\n+\t0x8a, 0x19, 0xdb, 0x31, 0xbb, 0x7b, 0x0f, 0x3b,\n+\t0x80, 0xa4, 0xbc, 0x2f, 0xaf, 0x72, 0x6b, 0xdb,\n+\t0x62, 0xaa, 0xba, 0xdd, 0xde, 0x77, 0xc6, 0x13,\n+\t0xeb, 0x9d, 0x0c, 0x78, 0x01, 0x01, 0xcd, 0x00,\n+\t0x32, 0xff, 0x74, 0x72, 0x65, 0x65, 0x20, 0x62,\n+\t0x36, 0x30, 0x39, 0x37, 0x37, 0x64, 0x37, 0x63,\n+\t0x34, 0x63, 0x32, 0x64, 0x31, 0x65, 0x63, 0x63,\n+\t0x33, 0x66, 0x62, 0x61, 0x31, 0x64, 0x39, 0x38,\n+\t0x65, 0x65, 0x31, 0x32, 0x30, 0x61, 0x64, 0x63,\n+\t0x32, 0x34, 0x38, 0x33, 0x34, 0x39, 0x35, 0x30,\n+\t0x62, 0x65, 0x34, 0x31, 0x32, 0x64, 0x39, 0x34,\n+\t0x63, 0x38, 0x30, 0x39, 0x34, 0x38, 0x30, 0x66,\n+\t0x35, 0x38, 0x62, 0x61, 0x39, 0x64, 0x61, 0x0a,\n+\t0x61, 0x75, 0x74, 0x68, 0x6f, 0x72, 0x20, 0x41,\n+\t0x20, 0x55, 0x20, 0x54, 0x68, 0x6f, 0x72, 0x20,\n+\t0x3c, 0x61, 0x75, 0x74, 0x68, 0x6f, 0x72, 0x40,\n+\t0x65, 0x78, 0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e,\n+\t0x63, 0x6f, 0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33,\n+\t0x34, 0x35, 0x36, 0x37, 0x38, 0x39, 0x30, 0x20,\n+\t0x2b, 0x30, 0x30, 0x30, 0x30, 0x0a, 0x63, 0x6f,\n+\t0x6d, 0x6d, 0x69, 0x74, 0x74, 0x65, 0x72, 0x20,\n+\t0x43, 0x20, 0x4f, 0x20, 0x4d, 0x69, 0x74, 0x74,\n+\t0x65, 0x72, 0x20, 0x3c, 0x63, 0x6f, 0x6d, 0x6d,\n+\t0x69, 0x74, 0x74, 0x65, 0x72, 0x40, 0x65, 0x78,\n+\t0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f,\n+\t0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33, 0x34, 0x35,\n+\t0x36, 0x37, 0x38, 0x39, 0x30, 0x20, 0x2b, 0x30,\n+\t0x30, 0x30, 0x30, 0x0a, 0x0a, 0x4c, 0x61, 0x72,\n+\t0x67, 0x65, 0x20, 0x62, 0x6c, 0x6f, 0x62, 0x20,\n+\t0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74, 0x0a, 0xb7,\n+\t0x80, 0x3d, 0xd7, 0x20, 0x78, 0x01, 0x01, 0x00,\n+\t0x00, 0xff, 0xff, 0x00, 0x00, 0x00, 0x01, 0x95,\n+\t0x11, 0x78, 0x01, 0x01, 0x15, 0x01, 0xea, 0xfe,\n+\t0x74, 0x72, 0x65, 0x65, 0x20, 0x36, 0x65, 0x66,\n+\t0x31, 0x39, 0x62, 0x34, 0x31, 0x32, 0x32, 0x35,\n+\t0x63, 0x35, 0x33, 0x36, 0x39, 0x66, 0x31, 0x63,\n+\t0x31, 0x30, 0x34, 0x64, 0x34, 0x35, 0x64, 0x38,\n+\t0x64, 0x38, 0x35, 0x65, 0x66, 0x61, 0x39, 0x62,\n+\t0x30, 0x35, 0x37, 0x62, 0x35, 0x33, 0x62, 0x31,\n+\t0x34, 0x62, 0x34, 0x62, 0x39, 0x62, 0x39, 0x33,\n+\t0x39, 0x64, 0x64, 0x37, 0x34, 0x64, 0x65, 0x63,\n+\t0x63, 0x35, 0x33, 0x32, 0x31, 0x0a, 0x70, 0x61,\n+\t0x72, 0x65, 0x6e, 0x74, 0x20, 0x37, 0x35, 0x62,\n+\t0x66, 0x30, 0x63, 0x34, 0x37, 0x61, 0x65, 0x34,\n+\t0x62, 0x62, 0x33, 0x30, 0x38, 0x65, 0x37, 0x63,\n+\t0x63, 0x32, 0x34, 0x38, 0x32, 0x65, 0x32, 0x32,\n+\t0x65, 0x66, 0x61, 0x65, 0x33, 0x37, 0x38, 0x37,\n+\t0x61, 0x39, 0x36, 0x38, 0x34, 0x38, 0x62, 0x64,\n+\t0x31, 0x37, 0x34, 0x39, 0x35, 0x36, 0x37, 0x31,\n+\t0x34, 0x37, 0x31, 0x35, 0x32, 0x34, 0x36, 0x64,\n+\t0x64, 0x62, 0x64, 0x35, 0x34, 0x0a, 0x61, 0x75,\n+\t0x74, 0x68, 0x6f, 0x72, 0x20, 0x41, 0x20, 0x55,\n+\t0x20, 0x54, 0x68, 0x6f, 0x72, 0x20, 0x3c, 0x61,\n+\t0x75, 0x74, 0x68, 0x6f, 0x72, 0x40, 0x65, 0x78,\n+\t0x61, 0x6d, 0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f,\n+\t0x6d, 0x3e, 0x20, 0x31, 0x32, 0x33, 0x34, 0x35,\n+\t0x36, 0x37, 0x38, 0x39, 0x30, 0x20, 0x2b, 0x30,\n+\t0x30, 0x30, 0x30, 0x0a, 0x63, 0x6f, 0x6d, 0x6d,\n+\t0x69, 0x74, 0x74, 0x65, 0x72, 0x20, 0x43, 0x20,\n+\t0x4f, 0x20, 0x4d, 0x69, 0x74, 0x74, 0x65, 0x72,\n+\t0x20, 0x3c, 0x63, 0x6f, 0x6d, 0x6d, 0x69, 0x74,\n+\t0x74, 0x65, 0x72, 0x40, 0x65, 0x78, 0x61, 0x6d,\n+\t0x70, 0x6c, 0x65, 0x2e, 0x63, 0x6f, 0x6d, 0x3e,\n+\t0x20, 0x31, 0x32, 0x33, 0x34, 0x35, 0x36, 0x37,\n+\t0x38, 0x39, 0x30, 0x20, 0x2b, 0x30, 0x30, 0x30,\n+\t0x30, 0x0a, 0x0a, 0x45, 0x6d, 0x70, 0x74, 0x79,\n+\t0x20, 0x74, 0x72, 0x65, 0x65, 0x20, 0x63, 0x6f,\n+\t0x6d, 0x6d, 0x69, 0x74, 0x0a, 0x6d, 0x6d, 0x51,\n+\t0x9a, 0xc9, 0x11, 0x76, 0x61, 0xa3, 0x89, 0x49,\n+\t0xb7, 0xa1, 0x58, 0xc6, 0x1d, 0x8c, 0x33, 0x75,\n+\t0x8d, 0x7e, 0x4d, 0x8e, 0x58, 0x91, 0xf8, 0x5c,\n+\t0x57, 0xd9, 0x89, 0x9e, 0xb8, 0xd2, 0x9a, 0xd8,\n+\t0xc9\n+};\n+\n static const struct fast_pack fast_packs[] = {\n \t{\n \t\t.format_id = GIT_SHA1_FORMAT_ID,\n@@ -253,6 +337,13 @@ static const struct fast_pack fast_packs[] = {\n \t\t.suffix_len = sizeof(fast_pack_sha1_suffix),\n \t\t.commit_oid = \"aac43daf40d0377af31aa9c798a4ae8a31b55c1d\",\n \t},\n+\t{\n+\t\t.format_id = GIT_SHA256_FORMAT_ID,\n+\t\t.suffix = fast_pack_sha256_suffix,\n+\t\t.suffix_len = sizeof(fast_pack_sha256_suffix),\n+\t\t.commit_oid = \"63c46ca51267b1d45be69a044bb84b4bf0559f09\"\n+\t\t\t      \"d727f861d2ae94ddebdddbc9\",\n+\t},\n };\n \n /*\n-- \ngitgitgadget\n\n"},{"id":"542899","messageId":"86c09af4f56deaa5ee91eec5ede5e640b46cdeb9.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 10/11] t5608: mark >4GB tests as EXPENSIVE","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:48Z","receivedAt":"2026-05-08T08:17:08Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nEven with precomputed pack constants that reduced the helper's\nruntime from minutes to seconds, the >4GB clone tests still take\n200-850 seconds across CI jobs. The bottleneck is no longer the\npack generation but the clone operations themselves: transporting,\nunpacking, and indexing 4 GiB of data through unpack-objects and\nindex-pack is inherently expensive.\n\nAs Jeff King pointed out [1], t5608 alone takes 160 seconds on his\nlaptop while the rest of the entire test suite finishes in under 90\nseconds, and the test's disk footprint (4+ GiB source repo, then\ntwo clones) is problematic for developers who use RAM disks for\ntheir trash directories.\n\nGate the >4GB tests on the EXPENSIVE prereq (which requires\nGIT_TEST_LONG to be set) in addition to SIZE_T_IS_64BIT, keeping\nthem out of normal local test runs.\n\n[1] https://lore.kernel.org/git/20260501063805.GA2038915@coredump.intra.peff.net/\n\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n t/t5608-clone-2gb.sh | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t5608-clone-2gb.sh b/t/t5608-clone-2gb.sh\nindex af93302dde..4f8a95ddda 100755\n--- a/t/t5608-clone-2gb.sh\n+++ b/t/t5608-clone-2gb.sh\n@@ -49,7 +49,7 @@ test_expect_success 'clone - with worktree, file:// protocol' '\n \n '\n \n-test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n+test_expect_success SIZE_T_IS_64BIT,EXPENSIVE 'set up repo with >4GB object' '\n \tlarge_blob_size=$((4*1024*1024*1024+1)) &&\n \tgit init --bare 4gb-repo &&\n \thead_oid=$(test-tool synthesize pack \\\n@@ -60,7 +60,7 @@ test_expect_success SIZE_T_IS_64BIT 'set up repo with >4GB object' '\n \tgit -C 4gb-repo symbolic-ref HEAD refs/heads/main\n '\n \n-test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n+test_expect_success SIZE_T_IS_64BIT,EXPENSIVE 'clone >4GB object via unpack-objects' '\n \t# The synthesized pack has five objects, so a large unpack limit keeps\n \t# fetch-pack on the unpack-objects path.\n \tgit -c fetch.unpackLimit=100 clone --bare \\\n@@ -76,7 +76,7 @@ test_expect_success SIZE_T_IS_64BIT 'clone >4GB object via unpack-objects' '\n \ttest \"$source_blob\" = \"$clone_blob\"\n '\n \n-test_expect_success SIZE_T_IS_64BIT 'clone with >4GB object via index-pack' '\n+test_expect_success SIZE_T_IS_64BIT,EXPENSIVE 'clone with >4GB object via index-pack' '\n \t# Force fetch-pack to hand the pack to index-pack instead.\n \tgit -c fetch.unpackLimit=1 clone --bare \\\n \t\t\"file://$(pwd)/4gb-repo\" 4gb-clone-index &&\n-- \ngitgitgadget\n\n"},{"id":"542900","messageId":"2159f6a271b06d156134392ce3c44fe957c83378.1778228209.git.gitgitgadget@gmail.com","threadId":"65563","inReplyTo":"pull.2102.v3.git.1778228209.gitgitgadget@gmail.com","subject":"[PATCH v3 11/11] ci: run expensive tests on push builds to integration branches","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-05-08T08:16:49Z","receivedAt":"2026-05-08T08:17:09Z","isPatch":true,"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nDerrick Stolee suggested [1] that expensive tests should be run at a\nregular cadence rather than on every PR iteration. Gate GIT_TEST_LONG\non push builds to the integration branches (next, master, main, maint)\nso that the EXPENSIVE prereq is satisfied there but not during PR\nvalidation, where the extra minutes of wall-clock time do not justify\nthemselves.\n\n[1] https://lore.kernel.org/git/e1e8837f-7374-4079-ba87-ab95dd156e33@gmail.com/\n\nHelped-by: Derrick Stolee <derrickstolee@github.com>\nAssisted-by: Claude Opus 4.6\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n ci/lib.sh | 9 +++++++++\n 1 file changed, 9 insertions(+)\n\ndiff --git a/ci/lib.sh b/ci/lib.sh\nindex 42a2b6a318..a671994bdf 100755\n--- a/ci/lib.sh\n+++ b/ci/lib.sh\n@@ -314,6 +314,15 @@ export DEFAULT_TEST_TARGET=prove\n export GIT_TEST_CLONE_2GB=true\n export SKIP_DASHED_BUILT_INS=YesPlease\n \n+# Enable expensive tests on push builds to integration branches, but\n+# not on PR builds where the extra time is not justified for every\n+# iteration.\n+case \"$GITHUB_EVENT_NAME,$CI_BRANCH\" in\n+push,*next*|push,*master*|push,*main*|push,*maint*)\n+\texport GIT_TEST_LONG=YesPlease\n+\t;;\n+esac\n+\n case \"$distro\" in\n ubuntu-*)\n \t# Python 2 is end of life, and Ubuntu 23.04 and newer don't actually\n-- \ngitgitgadget\n"},{"id":"542931","messageId":"20260508190947.GA25792@tb-raspi4","threadId":"65563","inReplyTo":"fa39d84b-ddbc-3943-5cca-078fb18db80d@gmx.de","subject":"Re: [PATCH v2 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-05-08T19:09:48Z","receivedAt":"2026-05-08T19:09:59Z","isPatch":true,"body":"On Fri, May 08, 2026 at 09:36:53AM +0200, Johannes Schindelin wrote:\n> Hi Torsten,\n> \n> On Tue, 5 May 2026, Torsten Bögershausen wrote:\n> \n> > On Mon, May 04, 2026 at 05:08:18PM +0000, Johannes Schindelin via GitGitGadget wrote:\n> > > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > > \n> > > [...]\n> > > @@ -524,7 +524,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n> > >  \t\t\t      struct object_id *oid)\n> > >  {\n> > >  \tunsigned char *p;\n> > > -\tunsigned long size, c;\n> > > +\tsize_t size;\n> > > +\tunsigned long c;\n> >\n> > Does this look a little bit strange ?\n> \n> Good point.\n> \n> > p points to an unsigned char (better would be *uint8_t)\n> > then it is dereferenced into an \"unsigned long\".\n> > Then it is masked with 0x7f\n> > In short: should \"c\" be declared as uint8_t ?\n> \n> Almost. It should be a `size_t`, so that we don't have to cast it when\n> shifting it. I'll include a fix in the next iteration.\n\nI think I was not very clear here.\nDe-referencing a long (or size_t) from any address is something\nI would try to avoid:\nSome processors do not like to read a 32 or 64 bit value from\nan uneven address (and throw an exeption).\nx86 processors just handle it, using more than one bus cycle.\n\nIn short: Please keep the up-cast and simply read an uint8_t.\n\n\n\n"},{"id":"542963","messageId":"xmqqik8w9eb8.fsf@gitster.g","threadId":"65563","inReplyTo":"20260508190947.GA25792@tb-raspi4","subject":"Re: [PATCH v2 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-10T02:41:15Z","receivedAt":"2026-05-10T02:41:17Z","isPatch":true,"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> On Fri, May 08, 2026 at 09:36:53AM +0200, Johannes Schindelin wrote:\n>> Hi Torsten,\n>> \n>> On Tue, 5 May 2026, Torsten Bögershausen wrote:\n>> \n>> > On Mon, May 04, 2026 at 05:08:18PM +0000, Johannes Schindelin via GitGitGadget wrote:\n>> > > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>> > > \n>> > > [...]\n>> > > @@ -524,7 +524,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n>> > >  \t\t\t      struct object_id *oid)\n>> > >  {\n>> > >  \tunsigned char *p;\n>> > > -\tunsigned long size, c;\n>> > > +\tsize_t size;\n>> > > +\tunsigned long c;\n>> >\n>> > Does this look a little bit strange ?\n>> \n>> Good point.\n>> \n>> > p points to an unsigned char (better would be *uint8_t)\n>> > then it is dereferenced into an \"unsigned long\".\n>> > Then it is masked with 0x7f\n>> > In short: should \"c\" be declared as uint8_t ?\n>> \n>> Almost. It should be a `size_t`, so that we don't have to cast it when\n>> shifting it. I'll include a fix in the next iteration.\n>\n> I think I was not very clear here.\n> De-referencing a long (or size_t) from any address is something\n> I would try to avoid:\n> Some processors do not like to read a 32 or 64 bit value from\n> an uneven address (and throw an exeption).\n> x86 processors just handle it, using more than one bus cycle.\n>\n> In short: Please keep the up-cast and simply read an uint8_t.\n\nHmph, I do not think there is \"up-cast\" to keep.  And we do not\ndereference a random pointer that would be suitable for unsigned\nchar * as if it were \"unsigned long *\" or \"size_t *\" in this code.\n\nThis came from commit 48fb7deb5bbd87933e7d314b73d7c1b52667f80f\n\nAuthor: Linus Torvalds <torvalds@linux-foundation.org>\nDate:   Wed Jun 17 17:22:27 2009 -0700\n\n    Fix big left-shifts of unsigned char\n    \n    Shifting 'unsigned char' or 'unsigned short' left can result in sign\n    extension errors, since the C integer promotion rules means that the\n    unsigned char/short will get implicitly promoted to a signed 'int' due to\n    the shift (or due to other operations).\n    \n    This normally doesn't matter, but if you shift things up sufficiently, it\n    will now set the sign bit in 'int', and a subsequent cast to a bigger type\n    (eg 'long' or 'unsigned long') will now sign-extend the value despite the\n    original expression being unsigned.\n    \n    One example of this would be something like\n    \n            unsigned long size;\n            unsigned char c;\n    \n            size += c << 24;\n    \n    where despite all the variables being unsigned, 'c << 24' ends up being a\n    signed entity, and will get sign-extended when then doing the addition in\n    an 'unsigned long' type.\n\n\nYou could rewrite Linus's example to\n\n\tunsigned char *cp;\n\tunsigned long size;\n\tunsigned char c;\n\n\tc = *cp;\n\tsize += ((unsigned long)c) << 24;\n\nWhile I am sympathetic to that position, I also would not mind\n\n\tunsigned char *cp;\n\tunsigned long size;\n\tunsigned long c;\n\n\tc = *cp;\n\tsize += c << 24;\n\nall that much.  In any case, such a \"clean-up\" has little to do with\nthe topic under discussion, and itshould be discussed separately on\nits own merit, most likely when the dust settles after this topic\nlands.  Let's not contaminate the patches that is \"a trivial rewrite\nthat is so obviously correct to fix the assumption that ulong and\nsize_t are of the same size everywhere\" with unrelated clean-up.\n\nThanks.\n\n"},{"id":"542967","messageId":"20260510091436.GA5880@tb-raspi4","threadId":"65563","inReplyTo":"xmqqik8w9eb8.fsf@gitster.g","subject":"Re: [PATCH v2 01/11] index-pack, unpack-objects: use size_t for object size","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-05-10T09:14:36Z","receivedAt":"2026-05-10T09:14:48Z","isPatch":true,"body":"> \n> Hmph, I do not think there is \"up-cast\" to keep.  And we do not\n> dereference a random pointer that would be suitable for unsigned\n> char * as if it were \"unsigned long *\" or \"size_t *\" in this code.\n> \n> This came from commit 48fb7deb5bbd87933e7d314b73d7c1b52667f80f\n> \n> Author: Linus Torvalds <torvalds@linux-foundation.org>\n> Date:   Wed Jun 17 17:22:27 2009 -0700\n> \n>     Fix big left-shifts of unsigned char\n>     \n>     Shifting 'unsigned char' or 'unsigned short' left can result in sign\n>     extension errors, since the C integer promotion rules means that the\n>     unsigned char/short will get implicitly promoted to a signed 'int' due to\n>     the shift (or due to other operations).\n>     \n>     This normally doesn't matter, but if you shift things up sufficiently, it\n>     will now set the sign bit in 'int', and a subsequent cast to a bigger type\n>     (eg 'long' or 'unsigned long') will now sign-extend the value despite the\n>     original expression being unsigned.\n>     \n>     One example of this would be something like\n>     \n>             unsigned long size;\n>             unsigned char c;\n>     \n>             size += c << 24;\n>     \n>     where despite all the variables being unsigned, 'c << 24' ends up being a\n>     signed entity, and will get sign-extended when then doing the addition in\n>     an 'unsigned long' type.\n> \n> \n> You could rewrite Linus's example to\n> \n> \tunsigned char *cp;\n> \tunsigned long size;\n> \tunsigned char c;\n> \n> \tc = *cp;\n> \tsize += ((unsigned long)c) << 24;\n> \n> While I am sympathetic to that position, I also would not mind\n> \n> \tunsigned char *cp;\n> \tunsigned long size;\n> \tunsigned long c;\n> \n> \tc = *cp;\n> \tsize += c << 24;\n> \n> all that much.  In any case, such a \"clean-up\" has little to do with\n> the topic under discussion, and itshould be discussed separately on\n> its own merit, most likely when the dust settles after this topic\n> lands.  Let's not contaminate the patches that is \"a trivial rewrite\n> that is so obviously correct to fix the assumption that ulong and\n> size_t are of the same size everywhere\" with unrelated clean-up.\n> \n> Thanks.\n> \n\nSorry for the confusion and noise.\nMy brain insisted to read\nc = *cp; // fetch 8 bits from memory, upcast to unsigned long\nas if we have written\nc = *(long*)cp; // fetch 32/64 bits from memory\nwhich is a completely different thing.\n\nIn short: all is good.\nThanks for digging and the patience.\n\n"},{"id":"542983","messageId":"xmqqjyta9630.fsf@gitster.g","threadId":"65563","inReplyTo":"2159f6a271b06d156134392ce3c44fe957c83378.1778228209.git.gitgitgadget@gmail.com","subject":"[PATCH] ci: enable EXPENSIVE for contributor builds","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-10T23:51:15Z","receivedAt":"2026-05-10T23:51:17Z","isPatch":true,"body":"Earlier, we enabled EXPENSIVE tests for pushes to integration\nbranches. As we didn't have any CI jobs that run these tests, this\nwas a step in the right direction.\n\nIt however is an ineffective and inefficient use of the maintainer\ntime, which does not scale, to allow contributors to send changes\nthat are less tested at the list, only to force the maintainer\nnotice breakages caused by their changes but only after these\nchanges are mixed with changes from other contributors.  The\nproblematic topic needs to be isolated by bisecting, and it\nhistorically has been done by the maintainer alone.\n\nIt is far better to let the problem identified early, preferably\nbefore the problematic code leaves the hands of the original\ndeveloper.  In order for it to happen, the test coverage of the\ncontributor tests must be at least as wide as the coverage of the\nintegration tests.\n\nEnable expensive tests for CI jobs triggered by pull requests.  This\nwill make each contributor take care of their own, which scales much\nbetter.\n\nKeep the expensive tests also enabled for the pushes of integration\nbranches, as that is the only place we can notice problems stemming\nfrom mismerges and inter-topic interactions, even if the topics from\nthe contributors in isolation all passes these tests.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n * This is to be applied on top of \"ci: run expensive tests on push\n   builds to integration branches\", currently sitting at the tip of\n   the js/objects-larger-than-4gb-on-windows topic.\n---\n ci/lib.sh | 10 ++++++----\n 1 file changed, 6 insertions(+), 4 deletions(-)\n\ndiff --git a/ci/lib.sh b/ci/lib.sh\nindex a671994bdf..4ca3ecef2c 100755\n--- a/ci/lib.sh\n+++ b/ci/lib.sh\n@@ -314,11 +314,13 @@ export DEFAULT_TEST_TARGET=prove\n export GIT_TEST_CLONE_2GB=true\n export SKIP_DASHED_BUILT_INS=YesPlease\n \n-# Enable expensive tests on push builds to integration branches, but\n-# not on PR builds where the extra time is not justified for every\n-# iteration.\n+# In order to give maximum test coverage to contributor builds,\n+# preferrably even before the changes consume public review bandwidth,\n+# enable \"expensive\" tests for PR events.\n+# In order to catch bugs introduced at integration time by mismerges,\n+# enable the long tests for pushes to the integration branches as well.\n case \"$GITHUB_EVENT_NAME,$CI_BRANCH\" in\n-push,*next*|push,*master*|push,*main*|push,*maint*)\n+pull_request,*|push,*next*|push,*master*|push,*main*|push,*maint*)\n \texport GIT_TEST_LONG=YesPlease\n \t;;\n esac\n-- \n2.54.0-162-gf1ca62098f\n\n"},{"id":"543023","messageId":"agF_0x0yq78J-RFk@pks.im","threadId":"65563","inReplyTo":"xmqqjyta9630.fsf@gitster.g","subject":"Re: [PATCH] ci: enable EXPENSIVE for contributor builds","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-05-11T07:05:55Z","receivedAt":"2026-05-11T07:06:03Z","isPatch":true,"body":"On Mon, May 11, 2026 at 08:51:15AM +0900, Junio C Hamano wrote:\n> diff --git a/ci/lib.sh b/ci/lib.sh\n> index a671994bdf..4ca3ecef2c 100755\n> --- a/ci/lib.sh\n> +++ b/ci/lib.sh\n> @@ -314,11 +314,13 @@ export DEFAULT_TEST_TARGET=prove\n>  export GIT_TEST_CLONE_2GB=true\n>  export SKIP_DASHED_BUILT_INS=YesPlease\n>  \n> -# Enable expensive tests on push builds to integration branches, but\n> -# not on PR builds where the extra time is not justified for every\n> -# iteration.\n> +# In order to give maximum test coverage to contributor builds,\n> +# preferrably even before the changes consume public review bandwidth,\n> +# enable \"expensive\" tests for PR events.\n> +# In order to catch bugs introduced at integration time by mismerges,\n> +# enable the long tests for pushes to the integration branches as well.\n>  case \"$GITHUB_EVENT_NAME,$CI_BRANCH\" in\n> -push,*next*|push,*master*|push,*main*|push,*maint*)\n> +pull_request,*|push,*next*|push,*master*|push,*main*|push,*maint*)\n>  \texport GIT_TEST_LONG=YesPlease\n>  \t;;\n>  esac\n\nSo with this change we now run the tests for all \"official\" branches,\nand on pull requests. Which raises the question: are there any events\nthat happen regularly that are excluded by this? Because if not I think\nit might be sensible to just enable this unconditionally, also because\nthat would make jobs on GitLab CI run expensive tests, as well.\n\nPatrick\n"},{"id":"543030","messageId":"xmqq33zys62a.fsf@gitster.g","threadId":"65563","inReplyTo":"agF_0x0yq78J-RFk@pks.im","subject":"Re: [PATCH] ci: enable EXPENSIVE for contributor builds","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-11T08:29:01Z","receivedAt":"2026-05-11T08:29:08Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> So with this change we now run the tests for all \"official\" branches,\n> and on pull requests. Which raises the question: are there any events\n> that happen regularly that are excluded by this? Because if not I think\n> it might be sensible to just enable this unconditionally, also because\n> that would make jobs on GitLab CI run expensive tests, as well.\n\nThe simplicity certainly is tempting.\n\nWe could instead do the \"let's enable only on linux-test-vars\" kind\nof \"optimization\", which is on the other side of the extreme, but\nthat is only valid if the kind of bugs that can be revealed only by\nEXPENSIVE tests, which may not be caught by others, is expected to\nbe pretty much platform or configuration agnostic.  I somehow doubt\nthat it is the case.\n\nIn any case, I think spending on more machine cycles is certainly\ncheaper than human resources for things like this.\n\nThanks.\n"},{"id":"543042","messageId":"agGpJlY_iTnzVoGr@pks.im","threadId":"65563","inReplyTo":"xmqq33zys62a.fsf@gitster.g","subject":"Re: [PATCH] ci: enable EXPENSIVE for contributor builds","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-05-11T10:02:14Z","receivedAt":"2026-05-11T10:02:20Z","isPatch":true,"body":"On Mon, May 11, 2026 at 05:29:01PM +0900, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > So with this change we now run the tests for all \"official\" branches,\n> > and on pull requests. Which raises the question: are there any events\n> > that happen regularly that are excluded by this? Because if not I think\n> > it might be sensible to just enable this unconditionally, also because\n> > that would make jobs on GitLab CI run expensive tests, as well.\n> \n> The simplicity certainly is tempting.\n> \n> We could instead do the \"let's enable only on linux-test-vars\" kind\n> of \"optimization\", which is on the other side of the extreme, but\n> that is only valid if the kind of bugs that can be revealed only by\n> EXPENSIVE tests, which may not be caught by others, is expected to\n> be pretty much platform or configuration agnostic.  I somehow doubt\n> that it is the case.\n\nYeah, it's probably not.\n\n> In any case, I think spending on more machine cycles is certainly\n> cheaper than human resources for things like this.\n\nAgreed.\n\nPatrick\n"}]}