{"thread":{"id":"65988","subject":"[PATCH 1/2] http: use unique tempfiles for packfile URI downloads","startedAt":"2026-07-13T22:34:39Z","lastAt":"2026-08-08T16:23:38Z","messageCount":57,"participants":["Ted Nyman","Junio C Hamano","Taylor Blau","Jeff King"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"548046","messageId":"alVn-QmK3K91_tkH@com-76773","threadId":"65988","inReplyTo":"cover.1783982021.git.tnyman@openai.com","subject":"[PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-13T22:34:33Z","receivedAt":"2026-07-13T22:34:39Z","isPatch":true,"body":"Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,\n2020-06-10), packfile URI downloads have been staged at\nobjects/pack/pack-<hash>.pack.temp.\n\nThe path is derived from the advertised pack hash. Two processes\nfetching the same pack into a shared object database therefore open the\nsame file for append. Their writes can corrupt the temporary pack. If\none process arrives after the other has completed the download, it may\ninstead try to resume at EOF, which some HTTP servers reject with 416.\n\nUse the tempfile API to give direct packfile URI downloads unique\ntemporary files. Keep the deterministic path for ordinary dumb HTTP\npack requests, which use it to resume a partial download left by an\nearlier invocation.\n\nThis means that a packfile URI download cannot be resumed by a later\ninvocation. A retry starts with an empty temporary file instead.\n\nAdd a test which pauses one process after downloading the pack and\nstarts another process using the same object database.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc |  5 +-\n http.c                            | 77 +++++++++++++++++++++----------\n http.h                            |  1 +\n t/t5550-http-fetch-dumb.sh        | 72 ++++++++++++++++++++++++++++-\n 4 files changed, 126 insertions(+), 29 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 2200f073c4..533bf381c4 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -48,9 +48,8 @@ commit-id::\n \tline (which is not expected in\n \tthis case), 'git http-fetch' fetches the packfile directly at the given\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n-\tThe hash is used to determine the name of the temporary file and is\n-\tarbitrary. The output of index-pack is printed to stdout. Requires\n-\t--index-pack-args.\n+\tThe hash is arbitrary. The output of index-pack is printed to stdout.\n+\tRequires --index-pack-args.\n \n --index-pack-args=<args>::\n \tFor internal use only. The command to run on the contents of the\ndiff --git a/http.c b/http.c\nindex b4e7b8d00b..5a46e7c65c 100644\n--- a/http.c\n+++ b/http.c\n@@ -2668,7 +2668,10 @@ int http_get_info_packs(const char *base_url, struct packfile_list *packs)\n \n void release_http_pack_request(struct http_pack_request *preq)\n {\n-\tif (preq->packfile) {\n+\tif (preq->tempfile) {\n+\t\tdelete_tempfile(&preq->tempfile);\n+\t\tpreq->packfile = NULL;\n+\t} else if (preq->packfile) {\n \t\tfclose(preq->packfile);\n \t\tpreq->packfile = NULL;\n \t}\n@@ -2688,7 +2691,10 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \tint tmpfile_fd;\n \tint ret = 0;\n \n-\tfclose(preq->packfile);\n+\tif (preq->tempfile)\n+\t\tclose_tempfile_gently(preq->tempfile);\n+\telse\n+\t\tfclose(preq->packfile);\n \tpreq->packfile = NULL;\n \n \ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n@@ -2711,7 +2717,10 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \n cleanup:\n \tclose(tmpfile_fd);\n-\tunlink(preq->tmpfile.buf);\n+\tif (preq->tempfile)\n+\t\tdelete_tempfile(&preq->tempfile);\n+\telse\n+\t\tunlink(preq->tmpfile.buf);\n \treturn ret;\n }\n \n@@ -2723,20 +2732,8 @@ void http_install_packfile(struct packed_git *p,\n \tpackfile_store_add_pack(files->packed, p);\n }\n \n-struct http_pack_request *new_http_pack_request(\n-\tconst unsigned char *packed_git_hash, const char *base_url) {\n-\n-\tstruct strbuf buf = STRBUF_INIT;\n-\n-\tend_url_with_slash(&buf, base_url);\n-\tstrbuf_addf(&buf, \"objects/pack/pack-%s.pack\",\n-\t\thash_to_hex(packed_git_hash));\n-\treturn new_direct_http_pack_request(packed_git_hash,\n-\t\t\t\t\t    strbuf_detach(&buf, NULL));\n-}\n-\n-struct http_pack_request *new_direct_http_pack_request(\n-\tconst unsigned char *packed_git_hash, char *url)\n+static struct http_pack_request *new_http_pack_request_for_url(\n+\tconst unsigned char *packed_git_hash, char *url, int resumable)\n {\n \toff_t prev_posn = 0;\n \tstruct http_pack_request *preq;\n@@ -2746,9 +2743,22 @@ struct http_pack_request *new_direct_http_pack_request(\n \n \tpreq->url = url;\n \n-\todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n-\tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n+\tif (resumable) {\n+\t\todb_pack_name(the_repository, &preq->tmpfile,\n+\t\t\t      packed_git_hash, \"pack\");\n+\t\tstrbuf_addstr(&preq->tmpfile, \".temp\");\n+\t\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n+\t} else {\n+\t\tstrbuf_addf(&preq->tmpfile, \"%s/pack/tmp_pack_XXXXXX\",\n+\t\t\t    repo_get_object_directory(the_repository));\n+\t\tpreq->tempfile = mks_tempfile_m(preq->tmpfile.buf, 0444);\n+\t\tif (preq->tempfile) {\n+\t\t\tstrbuf_reset(&preq->tmpfile);\n+\t\t\tstrbuf_addstr(&preq->tmpfile,\n+\t\t\t\t      get_tempfile_path(preq->tempfile));\n+\t\t\tpreq->packfile = fdopen_tempfile(preq->tempfile, \"w\");\n+\t\t}\n+\t}\n \tif (!preq->packfile) {\n \t\terror(\"Unable to open local file %s for pack\",\n \t\t      preq->tmpfile.buf);\n@@ -2766,8 +2776,9 @@ struct http_pack_request *new_direct_http_pack_request(\n \t * If there is data present from a previous transfer attempt,\n \t * resume where it left off\n \t */\n-\tprev_posn = ftello(preq->packfile);\n-\tif (prev_posn>0) {\n+\tif (resumable)\n+\t\tprev_posn = ftello(preq->packfile);\n+\tif (prev_posn > 0) {\n \t\tif (http_is_verbose)\n \t\t\tfprintf(stderr,\n \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\n@@ -2779,12 +2790,28 @@ struct http_pack_request *new_direct_http_pack_request(\n \treturn preq;\n \n abort:\n-\tstrbuf_release(&preq->tmpfile);\n-\tfree(preq->url);\n-\tfree(preq);\n+\trelease_http_pack_request(preq);\n \treturn NULL;\n }\n \n+struct http_pack_request *new_http_pack_request(\n+\tconst unsigned char *packed_git_hash, const char *base_url)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\n+\tend_url_with_slash(&buf, base_url);\n+\tstrbuf_addf(&buf, \"objects/pack/pack-%s.pack\",\n+\t\thash_to_hex(packed_git_hash));\n+\treturn new_http_pack_request_for_url(packed_git_hash,\n+\t\t\t\t\t     strbuf_detach(&buf, NULL), 1);\n+}\n+\n+struct http_pack_request *new_direct_http_pack_request(\n+\tconst unsigned char *packed_git_hash, char *url)\n+{\n+\treturn new_http_pack_request_for_url(packed_git_hash, url, 0);\n+}\n+\n /* Helpers for fetching objects (loose) */\n static size_t fwrite_sha1_file(char *ptr, size_t eltsize, size_t nmemb,\n \t\t\t       void *data)\ndiff --git a/http.h b/http.h\nindex 729c51904d..2c900779f5 100644\n--- a/http.h\n+++ b/http.h\n@@ -224,6 +224,7 @@ struct http_pack_request {\n \n \tFILE *packfile;\n \tstruct strbuf tmpfile;\n+\tstruct tempfile *tempfile;\n \tstruct active_request_slot *slot;\n \tstruct curl_slist *headers;\n };\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex b0080bf204..314a74c433 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -293,6 +293,74 @@ test_expect_success 'http-fetch --packfile' '\n \tgit -C packfileclient cat-file -e \"$HASH\"\n '\n \n+test_expect_success PIPE 'concurrent http-fetch --packfile' '\n+\tgit init packfileclient-concurrent &&\n+\tHASH=$(git -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git rev-parse HEAD) &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\tpackhash=$(basename \"$p\" .pack) &&\n+\tpackhash=${packhash#pack-} &&\n+\n+\tmkfifo first-ready first-continue &&\n+\texec 8<>first-ready &&\n+\texec 9<>first-continue &&\n+\twrite_script git-wait-index-pack <<-\\EOF &&\n+\techo ready >\"$GIT_TEST_WAIT_READY\" &&\n+\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n+\texec git index-pack \"$@\"\n+\tEOF\n+\n+\t# Hold the first download before it is indexed, so that the second\n+\t# download installs the pack first.\n+\t{\n+\t\t(\n+\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n+\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n+\t\t\tgit -C packfileclient-concurrent http-fetch \\\n+\t\t\t\t--packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=wait-index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin \\\n+\t\t\t\t--index-pack-arg=--keep \\\n+\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\techo continue >&9\n+\t\twait $first_pid 2>/dev/null || :\n+\t\texec 8>&-\n+\t\texec 9>&-\n+\t\trm -f first-ready first-continue git-wait-index-pack\n+\t\" &&\n+\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\tgit -C packfileclient-concurrent http-fetch \\\n+\t\t--packfile=\"$packhash\" \\\n+\t\t--index-pack-arg=index-pack \\\n+\t\t--index-pack-arg=--stdin \\\n+\t\t--index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n+\techo continue >&9 &&\n+\twait \"$first_pid\" &&\n+\n+\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect first.out &&\n+\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect second.out &&\n+\ttest_path_is_missing \\\n+\t\t\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tfind packfileclient-concurrent/.git/objects/pack \\\n+\t\t-name \"tmp_pack_*\" -print >tmpfiles &&\n+\ttest_must_be_empty tmpfiles &&\n+\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n@@ -313,7 +381,9 @@ test_expect_success 'http-fetch --packfile with corrupt pack' '\n \tgit init packfileclient &&\n \tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git && ls objects/pack/pack-*.pack) &&\n \ttest_must_fail git -C packfileclient http-fetch --packfile \\\n-\t\t\"$HTTPD_URL\"/dumb/repo_bad1.git/$p\n+\t\t\"$HTTPD_URL\"/dumb/repo_bad1.git/$p &&\n+\tfind packfileclient/.git/objects/pack -name \"tmp_pack_*\" -print >tmpfiles &&\n+\ttest_must_be_empty tmpfiles\n '\n \n test_expect_success 'fetch notices corrupt idx' '\n-- \n2.55.0\n\n"},{"id":"548047","messageId":"alVoA5-fDDPwKPZZ@com-76773","threadId":"65988","inReplyTo":"cover.1783982021.git.tnyman@openai.com","subject":"[PATCH 2/2] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-13T22:34:43Z","receivedAt":"2026-07-13T22:34:48Z","isPatch":true,"body":"When \"index-pack --keep\" creates a .keep file, it reports\n\"keep<TAB><hash>\". If the file already exists, index-pack leaves it\nuntouched and reports \"pack<TAB><hash>\" instead.\n\nSince dd4b732df7 (upload-pack: send part of packfile response as uri,\n2020-06-10), fetch-pack has accepted only the \"keep\" form for packs\ndownloaded through packfile URIs. A concurrent fetch can install the\nsame pack and create its .keep file before another process reaches\nindex-pack. The latter process then fails even though index-pack\ncompleted successfully.\n\nAccept both successful forms. Add a path to pack_lockfiles only for the\n\"keep\" form, so cleanup removes only a keep file created by the current\nprocess and preserves a pre-existing one.\n\nAdd a regression test which pre-creates a keep file and verifies that a\nfetch succeeds without changing it.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n fetch-pack.c           | 36 ++++++++++++++++++++----------------\n t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n 2 files changed, 51 insertions(+), 16 deletions(-)\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 120e01f3cf..a16b80177a 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1887,9 +1887,12 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t}\n \n \tfor (i = 0; i < packfile_uris.nr; i++) {\n+\t\tint created_keep = 0;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 1];\n+\t\tchar packname[GIT_MAX_HEXSZ + 6];\n+\t\tconst char *packhash;\n+\t\tconst int packname_len = the_hash_algo->hexsz + 6;\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1907,16 +1910,16 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n-\t\t    memcmp(packname, \"keep\\t\", 5))\n-\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n-\n-\t\tif (read_in_full(cmd.out, packname,\n-\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n-\t\t    packname[the_hash_algo->hexsz] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n-\n-\t\tpackname[the_hash_algo->hexsz] = '\\0';\n+\t\tif (read_in_full(cmd.out, packname, packname_len) != packname_len ||\n+\t\t    packname[packname_len - 1] != '\\n')\n+\t\t\tdie(\"fetch-pack: expected pack or keep, TAB, hash, \"\n+\t\t\t    \"then LF in http-fetch output\");\n+\t\tpackname[packname_len - 1] = '\\0';\n+\t\tif (skip_prefix(packname, \"keep\\t\", &packhash))\n+\t\t\tcreated_keep = 1;\n+\t\telse if (!skip_prefix(packname, \"pack\\t\", &packhash))\n+\t\t\tdie(\"fetch-pack: expected pack or keep, TAB, hash, \"\n+\t\t\t    \"then LF in http-fetch output\");\n \n \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n \n@@ -1925,16 +1928,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (finish_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to finish http-fetch\");\n \n-\t\tif (memcmp(packfile_uris.items[i].string, packname,\n+\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n \t\t\t   the_hash_algo->hexsz))\n \t\t\tdie(\"fetch-pack: pack downloaded from %s does not match expected hash %.*s\",\n \t\t\t    uri, (int) the_hash_algo->hexsz,\n \t\t\t    packfile_uris.items[i].string);\n \n-\t\tstring_list_append_nodup(pack_lockfiles,\n-\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n-\t\t\t\t\t\t repo_get_object_directory(the_repository),\n-\t\t\t\t\t\t packname));\n+\t\tif (created_keep)\n+\t\t\tstring_list_append_nodup(pack_lockfiles,\n+\t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n+\t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n+\t\t\t\t\t\t\t packhash));\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\ndiff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh\nindex 9f6cf4142d..1861eb7d7c 100755\n--- a/t/t5702-protocol-v2.sh\n+++ b/t/t5702-protocol-v2.sh\n@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '\n \t\tfetch \"$HTTPD_URL/smart/http_parent\"\n '\n \n+test_expect_success 'packfile URI preserves an existing keep file' '\n+\tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n+\trm -rf \"$P\" http_child keep.expect &&\n+\n+\tgit init \"$P\" &&\n+\tgit -C \"$P\" config uploadpack.allowsidebandall true &&\n+\n+\techo my-blob >\"$P/my-blob\" &&\n+\tgit -C \"$P\" add my-blob &&\n+\tgit -C \"$P\" commit -m x &&\n+\tconfigure_exclusion \"$P\" my-blob >h &&\n+\n+\tgit init http_child &&\n+\tpackhash=$(cat packh) &&\n+\tkeep=\"http_child/.git/objects/pack/pack-$packhash.keep\" &&\n+\techo pre-existing >\"$keep\" &&\n+\tcp \"$keep\" keep.expect &&\n+\n+\tGIT_TEST_SIDEBAND_ALL=1 \\\n+\tgit -C http_child -c protocol.version=2 \\\n+\t\t-c fetch.uriprotocols=http,https \\\n+\t\tfetch \"$HTTPD_URL/smart/http_parent\" &&\n+\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.pack\" &&\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.idx\" &&\n+\ttest_cmp keep.expect \"$keep\" &&\n+\tgit -C http_child cat-file -e \"$(cat h)\"\n+'\n+\n test_expect_success 'fetching with valid packfile URI but invalid hash fails' '\n \tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n \trm -rf \"$P\" http_child log &&\n-- \n2.55.0\n"},{"id":"548049","messageId":"cover.1783982021.git.tnyman@openai.com","threadId":"65988","inReplyTo":null,"subject":"[PATCH 0/2] packfile URIs: support concurrent downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-13T22:37:58Z","receivedAt":"2026-07-13T22:38:03Z","isPatch":true,"body":"Packfile URI downloads currently stage a pack at\nobjects/pack/pack-<hash>.pack.temp. Two Git processes fetching the same\npack into one object database can append to that file concurrently,\nwhich can corrupt the temporary pack or cause a resume request at EOF.\n\nThe first patch gives each direct packfile URI download a private\ntemporary file. Ordinary dumb HTTP pack requests retain their existing\nresumable staging behavior. A later packfile URI retry starts a new\ndownload.\n\nThe second patch handles the related .keep race. When another process\nhas already created the keep file, index-pack reports \"pack<TAB><hash>\"\ninstead of \"keep<TAB><hash>\". Accept both successful forms and remove\nonly keep files created by the current process.\n\nEach patch adds a regression test for its respective race.\n\nTed Nyman (2):\n  http: use unique tempfiles for packfile URI downloads\n  fetch-pack: accept \"pack\" output for packfile URIs\n\n Documentation/git-http-fetch.adoc |  5 +-\n fetch-pack.c                      | 36 ++++++++-------\n http.c                            | 77 +++++++++++++++++++++----------\n http.h                            |  1 +\n t/t5550-http-fetch-dumb.sh        | 72 ++++++++++++++++++++++++++++-\n t/t5702-protocol-v2.sh            | 31 +++++++++++++\n 6 files changed, 177 insertions(+), 45 deletions(-)\n\n\nbase-commit: e9019fcafe0040228b8631c30f97ae1adb61bcdc\n-- \n2.55.0\n"},{"id":"548052","messageId":"xmqqse5mv10a.fsf@gitster.g","threadId":"65988","inReplyTo":"alVn-QmK3K91_tkH@com-76773","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-14T01:00:21Z","receivedAt":"2026-07-14T01:00:24Z","isPatch":true,"body":"Ted Nyman <tnyman@openai.com> writes:\n\n> Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,\n> 2020-06-10), packfile URI downloads have been staged at\n> objects/pack/pack-<hash>.pack.temp.\n>\n> The path is derived from the advertised pack hash. Two processes\n> fetching the same pack into a shared object database therefore open the\n> same file for append. Their writes can corrupt the temporary pack. If\n> one process arrives after the other has completed the download, it may\n> instead try to resume at EOF, which some HTTP servers reject with 416.\n>\n> Use the tempfile API to give direct packfile URI downloads unique\n> temporary files. Keep the deterministic path for ordinary dumb HTTP\n> pack requests, which use it to resume a partial download left by an\n> earlier invocation.\n>\n> This means that a packfile URI download cannot be resumed by a later\n> invocation. A retry starts with an empty temporary file instead.\n\nWhile that does sound like a safe and correct approach, stepping\nback briefly, would it not be wasteful for the second process to\ndownload the same packfile that the first has already started\ndownloading?\n\nAre there better ways for these processes to coordinate with each\nother?  Instead of appending to the file, what if the second process\nuses a predictable temporary name (which we already use) to open a\nnew file with O_CREAT | O_EXCL to avoid this redundant work?  If the\nopen call fails because the file already exists, the second process\ncan detect that another process is active and wait for it to finish\nrather than initiating its own network request.\n\nDoing so might require setting up a trigger or polling mechanism to\nwait for the first process's download to complete (and detecting if\nthe other process dies without cleaning up), though that may open a\ncan of worms.\n\n"},{"id":"548053","messageId":"alWXwAGWgXSXoRJv@com-76773","threadId":"65988","inReplyTo":"xmqqse5mv10a.fsf@gitster.g","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-14T01:58:24Z","receivedAt":"2026-07-14T01:58:30Z","isPatch":true,"body":"> While that does sound like a safe and correct approach, stepping\n> back briefly, would it not be wasteful for the second process to\n> download the same packfile that the first has already started\n> downloading?\n\nYes. If two fetches overlap, the second download is redundant.\n\n> Are there better ways for these processes to coordinate with each\n> other? Instead of appending to the file, what if the second process\n> uses a predictable temporary name (which we already use) to open a\n> new file with O_CREAT | O_EXCL to avoid this redundant work?\n\nUsing the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL\nwould prevent concurrent writes, but EEXIST alone would not\ndistinguish an in-progress download from one left by an earlier\nfailed or interrupted invocation. The existing .pack.temp name is not\ncovered by the tmp_* pruning path, so simply waiting for it to\ndisappear could leave a fetch stuck after a crash.\n\nThe waiting case would also need a complete handoff. If the first\nprocess finishes, the second would need to notice the installed pack\nand account for the expected index-pack result and keep state. If the\nfirst process fails and removes its temporary file, the second would\nneed to retry as the downloader. That is possible, but introduces\ncross-process coordination and a timeout policy in http-fetch.\n\nThe unique tempfile preserves the existing \"download, index, then\ninstall\" behavior for each invocation and fixes both the\nconcurrent-append and EOF-resume failures. Avoiding the duplicate\ntransfer would be useful for large packs, but I would prefer to keep\nthat as a follow-up unless you think it is necessary for this\ncorrectness fix.\n\nThanks,\nTed\n"},{"id":"548072","messageId":"alW1tAnMtOznxrhK@com-79390","threadId":"65988","inReplyTo":"alVn-QmK3K91_tkH@com-76773","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Taylor Blau","fromEmail":"ttaylorr@openai.com","sentAt":"2026-07-14T04:06:12Z","receivedAt":"2026-07-14T04:06:17Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 03:34:33PM -0700, Ted Nyman wrote:\n> Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,\n> 2020-06-10), packfile URI downloads have been staged at\n> objects/pack/pack-<hash>.pack.temp.\n>\n> The path is derived from the advertised pack hash. Two processes\n> fetching the same pack into a shared object database therefore open the\n> same file for append. Their writes can corrupt the temporary pack. If\n> one process arrives after the other has completed the download, it may\n> instead try to resume at EOF, which some HTTP servers reject with 416.\n>\n> Use the tempfile API to give direct packfile URI downloads unique\n> temporary files. Keep the deterministic path for ordinary dumb HTTP\n> pack requests, which use it to resume a partial download left by an\n> earlier invocation.\n>\n> This means that a packfile URI download cannot be resumed by a later\n> invocation. A retry starts with an empty temporary file instead.\n>\n> Add a test which pauses one process after downloading the pack and\n> starts another process using the same object database.\n>\n> Signed-off-by: Ted Nyman <tnyman@openai.com>\n> ---\n>  Documentation/git-http-fetch.adoc |  5 +-\n>  http.c                            | 77 +++++++++++++++++++++----------\n>  http.h                            |  1 +\n>  t/t5550-http-fetch-dumb.sh        | 72 ++++++++++++++++++++++++++++-\n>  4 files changed, 126 insertions(+), 29 deletions(-)\n>\n> diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\n> index 2200f073c4..533bf381c4 100644\n> --- a/Documentation/git-http-fetch.adoc\n> +++ b/Documentation/git-http-fetch.adoc\n> @@ -48,9 +48,8 @@ commit-id::\n>  \tline (which is not expected in\n>  \tthis case), 'git http-fetch' fetches the packfile directly at the given\n>  \tURL and uses index-pack to generate corresponding .idx and .keep files.\n> -\tThe hash is used to determine the name of the temporary file and is\n> -\tarbitrary. The output of index-pack is printed to stdout. Requires\n> -\t--index-pack-args.\n> +\tThe hash is arbitrary. The output of index-pack is printed to stdout.\n> +\tRequires --index-pack-args.\n>\n>  --index-pack-args=<args>::\n>  \tFor internal use only. The command to run on the contents of the\n> diff --git a/http.c b/http.c\n> index b4e7b8d00b..5a46e7c65c 100644\n> --- a/http.c\n> +++ b/http.c\n> @@ -2668,7 +2668,10 @@ int http_get_info_packs(const char *base_url, struct packfile_list *packs)\n>\n>  void release_http_pack_request(struct http_pack_request *preq)\n>  {\n> -\tif (preq->packfile) {\n> +\tif (preq->tempfile) {\n> +\t\tdelete_tempfile(&preq->tempfile);\n> +\t\tpreq->packfile = NULL;\n\nWe should be able to drop the assignment to NULL on the second line,\nsince `delete_tempfile()` takes a double pointer to the 'struct\npackfile' and NULL's it out for us.\n\n(The other callers appear to avoid explicitly setting `preq->tempfile`\nto NULL.)\n\nThe rest of the patch looks good to me.\n\n> diff --git a/http.h b/http.h\n> index 729c51904d..2c900779f5 100644\n> --- a/http.h\n> +++ b/http.h\n> @@ -224,6 +224,7 @@ struct http_pack_request {\n>\n>  \tFILE *packfile;\n>  \tstruct strbuf tmpfile;\n> +\tstruct tempfile *tempfile;\n>  \tstruct active_request_slot *slot;\n>  \tstruct curl_slist *headers;\n>  };\n> diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\n> index b0080bf204..314a74c433 100755\n> --- a/t/t5550-http-fetch-dumb.sh\n> +++ b/t/t5550-http-fetch-dumb.sh\n> @@ -293,6 +293,74 @@ test_expect_success 'http-fetch --packfile' '\n>  \tgit -C packfileclient cat-file -e \"$HASH\"\n>  '\n>\n> +test_expect_success PIPE 'concurrent http-fetch --packfile' '\n\nPhew ;-).\n\nThis is definitely tricky to test, but what you wrote here looks\nplausibly correct to me.\n\nThanks,\nTaylor\n"},{"id":"548073","messageId":"alW2EnNR21VmkESW@com-79390","threadId":"65988","inReplyTo":"alWXwAGWgXSXoRJv@com-76773","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Taylor Blau","fromEmail":"ttaylorr@openai.com","sentAt":"2026-07-14T04:07:46Z","receivedAt":"2026-07-14T04:07:50Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 06:58:24PM -0700, Ted Nyman wrote:\n> > While that does sound like a safe and correct approach, stepping\n> > back briefly, would it not be wasteful for the second process to\n> > download the same packfile that the first has already started\n> > downloading?\n>\n> Yes. If two fetches overlap, the second download is redundant.\n>\n> > Are there better ways for these processes to coordinate with each\n> > other? Instead of appending to the file, what if the second process\n> > uses a predictable temporary name (which we already use) to open a\n> > new file with O_CREAT | O_EXCL to avoid this redundant work?\n>\n> Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL\n> would prevent concurrent writes, but EEXIST alone would not\n> distinguish an in-progress download from one left by an earlier\n> failed or interrupted invocation. The existing .pack.temp name is not\n> covered by the tmp_* pruning path, so simply waiting for it to\n> disappear could leave a fetch stuck after a crash.\n\nExactly. If two processes are downloading the same pack at the same time\nto different locations, the effort is of course redundant. But I don't\nthink we can reliably distinguish between that case and one where an\nearlier process died in the middle of downloading a pack but was unable\nto clean up after itself.\n\nThanks,\nTaylor\n"},{"id":"548074","messageId":"alW3Wm2scg8TjPXy@com-79390","threadId":"65988","inReplyTo":"cover.1783982021.git.tnyman@openai.com","subject":"Re: [PATCH 0/2] packfile URIs: support concurrent downloads","fromName":"Taylor Blau","fromEmail":"ttaylorr@openai.com","sentAt":"2026-07-14T04:13:14Z","receivedAt":"2026-07-14T04:13:19Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 03:37:58PM -0700, Ted Nyman wrote:\n> Ted Nyman (2):\n>   http: use unique tempfiles for packfile URI downloads\n>   fetch-pack: accept \"pack\" output for packfile URIs\n\nI left one pretty minor style-nit on the first patch, but otherwise this\nlooks good to me.\n\nThanks,\nTaylor\n"},{"id":"548076","messageId":"20260714052833.GA2516582@coredump.intra.peff.net","threadId":"65988","inReplyTo":"alWXwAGWgXSXoRJv@com-76773","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-14T05:28:33Z","receivedAt":"2026-07-14T05:28:40Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 06:58:24PM -0700, Ted Nyman wrote:\n\n> > Are there better ways for these processes to coordinate with each\n> > other? Instead of appending to the file, what if the second process\n> > uses a predictable temporary name (which we already use) to open a\n> > new file with O_CREAT | O_EXCL to avoid this redundant work?\n> \n> Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL\n> would prevent concurrent writes, but EEXIST alone would not\n> distinguish an in-progress download from one left by an earlier\n> failed or interrupted invocation. The existing .pack.temp name is not\n> covered by the tmp_* pruning path, so simply waiting for it to\n> disappear could leave a fetch stuck after a crash.\n\nA few thoughts:\n\n  - Using O_EXCL makes this essentially a lockfile. So we could apply\n    the logic used elsewhere for lockfiles, like auto-removing files\n    with ancient mtimes. Or we could even go all-in with a pid check for\n    liveness; most of Git's lockfiles don't do that, but at least one\n    does (the background auto-gc lock).\n\n  - If we're not already using a name which is auto-cleaned during\n    maintenance, we probably ought to be. Leaving aside concurrency\n    issues, nobody would ever clean up the on-disk cruft.\n\n    But of course the original code here is intentionally _not_ using a\n    name we'd clean up, because it wants to be able to resume an\n    interrupted transfer.  And you're explicitly breaking that for the\n    packfile URI case.\n\n    Is that a cost we're OK with paying? Fixing it opens up that same\n    coordination can of worms. You have to tell the difference a\n    concurrent writer and a previous dead one (whose work you can\n    resume).\n\n    It does feel weird that we'd do one thing for dumb-http and another\n    for packfile URIs. Wouldn't they suffer from the same concurrency\n    and resumption problems?\n\n> The unique tempfile preserves the existing \"download, index, then\n> install\" behavior for each invocation and fixes both the\n> concurrent-append and EOF-resume failures. Avoiding the duplicate\n> transfer would be useful for large packs, but I would prefer to keep\n> that as a follow-up unless you think it is necessary for this\n> correctness fix.\n\nIf we're OK with killing the ability to resume, then yeah, I think it\nwould make sense to start simple and un-break things. And then put a\ncoordination layer on top later (or never if nobody cares enough).\n\n-Peff\n"},{"id":"548079","messageId":"20260714054439.GB2516582@coredump.intra.peff.net","threadId":"65988","inReplyTo":"alW1tAnMtOznxrhK@com-79390","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-14T05:44:39Z","receivedAt":"2026-07-14T05:44:41Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 09:06:12PM -0700, Taylor Blau wrote:\n\n> >  void release_http_pack_request(struct http_pack_request *preq)\n> >  {\n> > -\tif (preq->packfile) {\n> > +\tif (preq->tempfile) {\n> > +\t\tdelete_tempfile(&preq->tempfile);\n> > +\t\tpreq->packfile = NULL;\n> \n> We should be able to drop the assignment to NULL on the second line,\n> since `delete_tempfile()` takes a double pointer to the 'struct\n> packfile' and NULL's it out for us.\n> \n> (The other callers appear to avoid explicitly setting `preq->tempfile`\n> to NULL.)\n\nIt takes a double-pointer to the \"struct tempfile\"; the NULL assignment\nis to the \"packfile\" member, which is the FILE handle.\n\nI thought at first this was buggy; we still call fdopen() on the\ntempfile and assign the result to preq->packfile, even in the new\nnon-resumable case. Don't we need to fclose() it? But the answer is no:\nfdopen_tempfile() retains ownership of the result, storing it in\ntempfile.fp. So it will be correctly closed during delete_tempfile(),\nand in fact we must _not_ fclose it again.\n\nBut assigning NULL can happen with either style. So doing it\nunconditionally like:\n\n  if (preq->tempfile)\n\tdelete_tempfile(&preq->tempfile);\n  else if (preq->packfile)\n\tfclose(preq->packfile);\n  preq->packfile = NULL;\n\nmakes more sense, as it is done in finish_http_pack_request(). It might\neven make sense to add a comment explaining why we don't need to\nfclose() in the first part of the conditional.\n\nAll that said, I do not think setting it to NULL matters at all here,\nsince the function ends with free(preq). So just dropping the NULL would\nperhaps be more clear.\n\n-Peff\n"},{"id":"548080","messageId":"20260714064619.GC2516582@coredump.intra.peff.net","threadId":"65988","inReplyTo":"alVn-QmK3K91_tkH@com-76773","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-14T06:46:19Z","receivedAt":"2026-07-14T06:46:21Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 03:34:33PM -0700, Ted Nyman wrote:\n\n> The path is derived from the advertised pack hash. Two processes\n> fetching the same pack into a shared object database therefore open the\n> same file for append. Their writes can corrupt the temporary pack. If\n> one process arrives after the other has completed the download, it may\n> instead try to resume at EOF, which some HTTP servers reject with 416.\n\nYuck. In theory they're writing the same thing, but I think the source\nof the corruption is append mode. Two concurrent writers will keep\nauto-seeking to the end of the file, rather than keeping their own file\npointers. There's no way to ask for O_APPEND without O_TRUNC via stdio,\nbut we can drop down a level like this:\n\ndiff --git a/http.c b/http.c\nindex b4e7b8d00b..d7362c99a2 100644\n--- a/http.c\n+++ b/http.c\n@@ -2740,6 +2740,7 @@ struct http_pack_request *new_direct_http_pack_request(\n {\n \toff_t prev_posn = 0;\n \tstruct http_pack_request *preq;\n+\tint fd;\n \n \tCALLOC_ARRAY(preq, 1);\n \tstrbuf_init(&preq->tmpfile, 0);\n@@ -2748,12 +2749,13 @@ struct http_pack_request *new_direct_http_pack_request(\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n-\tif (!preq->packfile) {\n+\tfd = open(preq->tmpfile.buf, O_WRONLY|O_CREAT, 0666);\n+\tif (fd < 0) {\n \t\terror(\"Unable to open local file %s for pack\",\n \t\t      preq->tmpfile.buf);\n \t\tgoto abort;\n \t}\n+\tpreq->packfile = xfdopen(fd, \"w\");\n \n \tpreq->slot = get_active_slot();\n \tpreq->headers = object_request_headers();\n\nThat patch (with no other code changes) passes your test.\n\nI suspect it could cause us to racily send an http range of \"N-\" to the\nserver, where N is the total number of bytes in the file (because we\ndon't know how many bytes there are supposed to be). I don't know if\nthat would cause an HTTP 416 or not. I think possibly not, and the 416\nyou saw (and that I see when running the test without any code changes)\nmight be from sending a range that starts _past_ N. We end up with a\ntoo-long when both processes are appending.\n\nI can't say I love the overall notion of \"two processes are writing the\nsame data, it will probably be fine!\". There might be portability\nissues, and I'm not sure what would happen if we ever did get\nconflicting data. If we're just feeding this to \"index-pack --stdin\"\nwe'd at least notice the problem (rather than quietly corrupting the\nindexed file!).\n\nSo I'm offering this as a point for further discussion, and not\nnecessarily a counter-proposal. ;)\n\n> Use the tempfile API to give direct packfile URI downloads unique\n> temporary files. Keep the deterministic path for ordinary dumb HTTP\n> pack requests, which use it to resume a partial download left by an\n> earlier invocation.\n> \n> This means that a packfile URI download cannot be resumed by a later\n> invocation. A retry starts with an empty temporary file instead.\n\nArguably losing the ability to retry is a regression. In general, I\nthink we should prefer correctness to efficiency. But I wonder if this\nis a case where the user might want to make the choice to say \"I am not\ngoing to fetch two packfiles at once; please enable resumable fetches\".\nEspecially because one of the selling points of packfile URIs is that\nthey are resumable.\n\nOne other thought on resumable transfers: if we are not going to resume\nthe transfer, then why spool the pack to disk at all? In other words,\nwhy not just send it straight to \"index-pack --stdin\". That fixes your\nconcurrency issue (because it uses its own tempfiles behind the scene),\nbut has two other big advantages:\n\n  1. It halves the number of disk writes, and lowers the peak disk usage\n     (with the current code, there is a moment where both the tempfile\n     and the indexed pack are present on disk).\n\n  2. It pipelines the data processing. The current code bottlenecks on\n     the network while the CPU sits idle, and then bottlenecks on the\n     CPU once we have the whole file. We could be doing useful CPU work\n     during the network transfer, just like a regular pack code does.\n\n\nSo I'm not quite sold on losing the ability to resume entirely. And in\ncases where we do lose it, I think it opens up other improvements.\n\nBut I'll reader over the rest of the patch with the notion that this is\nthe direction we want to go in.\n\n> diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\n> index 2200f073c4..533bf381c4 100644\n> --- a/Documentation/git-http-fetch.adoc\n> +++ b/Documentation/git-http-fetch.adoc\n> @@ -48,9 +48,8 @@ commit-id::\n>  \tline (which is not expected in\n>  \tthis case), 'git http-fetch' fetches the packfile directly at the given\n>  \tURL and uses index-pack to generate corresponding .idx and .keep files.\n> -\tThe hash is used to determine the name of the temporary file and is\n> -\tarbitrary. The output of index-pack is printed to stdout. Requires\n> -\t--index-pack-args.\n> +\tThe hash is arbitrary. The output of index-pack is printed to stdout.\n> +\tRequires --index-pack-args.\n\nDo we even need to provide a hash anymore? After your patch I don't\nthink we even use it. It might be worth keeping around, though, as it\nwould be a unique key for de-duping or resuming, if we ever did\nimplement those on top.\n\n>  void release_http_pack_request(struct http_pack_request *preq)\n>  {\n> -\tif (preq->packfile) {\n> +\tif (preq->tempfile) {\n> +\t\tdelete_tempfile(&preq->tempfile);\n> +\t\tpreq->packfile = NULL;\n> +\t} else if (preq->packfile) {\n>  \t\tfclose(preq->packfile);\n>  \t\tpreq->packfile = NULL;\n>  \t}\n\nOK. I think this is correct, though see my comments elsewhere in the\nthread.\n\n> @@ -2688,7 +2691,10 @@ int finish_http_pack_request(struct http_pack_request *preq)\n>  \tint tmpfile_fd;\n>  \tint ret = 0;\n>  \n> -\tfclose(preq->packfile);\n> +\tif (preq->tempfile)\n> +\t\tclose_tempfile_gently(preq->tempfile);\n> +\telse\n> +\t\tfclose(preq->packfile);\n>  \tpreq->packfile = NULL;\n\nOK, and this is correct because preq->packfile is just an alias for\npreq->tempfile.fp when the tempfile is valid. The NULL assignment is\nimportant here so that the release() function doesn't double-free.\n\n> -struct http_pack_request *new_http_pack_request(\n> -\tconst unsigned char *packed_git_hash, const char *base_url) {\n> -\n> -\tstruct strbuf buf = STRBUF_INIT;\n> -\n> -\tend_url_with_slash(&buf, base_url);\n> -\tstrbuf_addf(&buf, \"objects/pack/pack-%s.pack\",\n> -\t\thash_to_hex(packed_git_hash));\n> -\treturn new_direct_http_pack_request(packed_git_hash,\n> -\t\t\t\t\t    strbuf_detach(&buf, NULL));\n> -}\n\nThis hunk puzzled me at first, but it's because we used to just be a\nwrapper for the \"direct\" variant, and now the two will share a single\nstatic helper. That might have been a little more clear as a preparatory\npatch, but OK.\n\n> +\tif (resumable) {\n> +\t\todb_pack_name(the_repository, &preq->tmpfile,\n> +\t\t\t      packed_git_hash, \"pack\");\n> +\t\tstrbuf_addstr(&preq->tmpfile, \".temp\");\n> +\t\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n> +\t} else {\n> +\t\tstrbuf_addf(&preq->tmpfile, \"%s/pack/tmp_pack_XXXXXX\",\n> +\t\t\t    repo_get_object_directory(the_repository));\n> +\t\tpreq->tempfile = mks_tempfile_m(preq->tmpfile.buf, 0444);\n> +\t\tif (preq->tempfile) {\n> +\t\t\tstrbuf_reset(&preq->tmpfile);\n> +\t\t\tstrbuf_addstr(&preq->tmpfile,\n> +\t\t\t\t      get_tempfile_path(preq->tempfile));\n> +\t\t\tpreq->packfile = fdopen_tempfile(preq->tempfile, \"w\");\n> +\t\t}\n> +\t}\n>  \tif (!preq->packfile) {\n>  \t\terror(\"Unable to open local file %s for pack\",\n>  \t\t      preq->tmpfile.buf);\n\nOK, and this is the meat of the change. We usually use odb_mkstemp() for\ntmp_pack_* files, but that annoyingly doesn't give you a tempfile\nstruct. So setting up your own filename and using mks_tempfile_m() makes\nsense here.\n\nThe error path is a little funny, but we catch it in the context when\npreq->packfile is NULL. Good.\n\n> @@ -2766,8 +2776,9 @@ struct http_pack_request *new_direct_http_pack_request(\n>  \t * If there is data present from a previous transfer attempt,\n>  \t * resume where it left off\n>  \t */\n> -\tprev_posn = ftello(preq->packfile);\n> -\tif (prev_posn>0) {\n> +\tif (resumable)\n> +\t\tprev_posn = ftello(preq->packfile);\n> +\tif (prev_posn > 0) {\n\nI think this is not technically necessary, as ftello() would just return\n\"0\" for our newly-created file. But it does make the intent clear.\n\n> @@ -2779,12 +2790,28 @@ struct http_pack_request *new_direct_http_pack_request(\n>  \treturn preq;\n>  \n>  abort:\n> -\tstrbuf_release(&preq->tmpfile);\n> -\tfree(preq->url);\n> -\tfree(preq);\n> +\trelease_http_pack_request(preq);\n>  \treturn NULL;\n>  }\n\nOK, now we have potentially more to free, so we rely on the release\nfunction. That could cause problems if we jump to this abort label when\nthe struct isn't fully initialized. I think it is OK, though. We zero\nthe whole thing, so the extra fields that the release() function\nconsiders will just be ignored.\n\n> diff --git a/http.h b/http.h\n> index 729c51904d..2c900779f5 100644\n> --- a/http.h\n> +++ b/http.h\n> @@ -224,6 +224,7 @@ struct http_pack_request {\n>  \n>  \tFILE *packfile;\n>  \tstruct strbuf tmpfile;\n> +\tstruct tempfile *tempfile;\n>  \tstruct active_request_slot *slot;\n>  \tstruct curl_slist *headers;\n\nYuck, now we have \"tempfile\" and \"tmpfile\" with two different types and\ntotally different semantics (and even when \"tempfile\" is in use,\n\"tmpfile\" is still meaningful!).\n\nCan we even just call the second one non_resumable_tempfile or\nsomething? It's a mouthful, but it makes it less likely to confuse the\ntwo.\n\n> +\t# Hold the first download before it is indexed, so that the second\n> +\t# download installs the pack first.\n> +\t{\n> +\t\t(\n> +\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n> +\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n> +\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n> +\t\t\tgit -C packfileclient-concurrent http-fetch \\\n> +\t\t\t\t--packfile=\"$packhash\" \\\n> +\t\t\t\t--index-pack-arg=wait-index-pack \\\n> +\t\t\t\t--index-pack-arg=--stdin \\\n> +\t\t\t\t--index-pack-arg=--keep \\\n> +\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n> +\t\t\tthen\n> +\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n> +\t\t\t\texit 1\n> +\t\t\tfi\n> +\t\t) &\n> +\t\tfirst_pid=$!\n> +\t} &&\n\nOK. I wonder if it would be simpler and a more robust test if rather\nthan writing the correct bytes (and then waiting), the first process\njust wrote total garbage. Then we'd be sure the other process is not\nreading it, because it would definitely corrupt their input.\n\nI dunno. This is a more realistic scenario, so in that sense maybe it is\nmore interesting.\n\n> +\ttest_when_finished \"\n> +\t\techo continue >&9\n> +\t\twait $first_pid 2>/dev/null || :\n> +\t\texec 8>&-\n> +\t\texec 9>&-\n> +\t\trm -f first-ready first-continue git-wait-index-pack\n> +\t\" &&\n> [...]\n\nThe rest of the fifo handling looks plausibly correct. This is a tricky\narea and it's common to introduce funky races, but I didn't see anything\nwrong, and it passed a few dozen rounds of --stress.\n\n> @@ -313,7 +381,9 @@ test_expect_success 'http-fetch --packfile with corrupt pack' '\n>  \tgit init packfileclient &&\n>  \tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git && ls objects/pack/pack-*.pack) &&\n>  \ttest_must_fail git -C packfileclient http-fetch --packfile \\\n> -\t\t\"$HTTPD_URL\"/dumb/repo_bad1.git/$p\n> +\t\t\"$HTTPD_URL\"/dumb/repo_bad1.git/$p &&\n> +\tfind packfileclient/.git/objects/pack -name \"tmp_pack_*\" -print >tmpfiles &&\n> +\ttest_must_be_empty tmpfiles\n>  '\n\nOK, so here we just detect that we cleaned up after ourselves. Makes\nsense.\n\n-Peff\n"},{"id":"548081","messageId":"20260714071231.GD2516582@coredump.intra.peff.net","threadId":"65988","inReplyTo":"alVoA5-fDDPwKPZZ@com-76773","subject":"Re: [PATCH 2/2] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-14T07:12:31Z","receivedAt":"2026-07-14T07:12:34Z","isPatch":true,"body":"On Mon, Jul 13, 2026 at 03:34:43PM -0700, Ted Nyman wrote:\n\n> When \"index-pack --keep\" creates a .keep file, it reports\n> \"keep<TAB><hash>\". If the file already exists, index-pack leaves it\n> untouched and reports \"pack<TAB><hash>\" instead.\n> \n> Since dd4b732df7 (upload-pack: send part of packfile response as uri,\n> 2020-06-10), fetch-pack has accepted only the \"keep\" form for packs\n> downloaded through packfile URIs. A concurrent fetch can install the\n> same pack and create its .keep file before another process reaches\n> index-pack. The latter process then fails even though index-pack\n> completed successfully.\n> \n> Accept both successful forms. Add a path to pack_lockfiles only for the\n> \"keep\" form, so cleanup removes only a keep file created by the current\n> process and preserves a pre-existing one.\n\nOK, that all makes sense.\n\n>  \tfor (i = 0; i < packfile_uris.nr; i++) {\n> +\t\tint created_keep = 0;\n>  \t\tint j;\n>  \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n> -\t\tchar packname[GIT_MAX_HEXSZ + 1];\n> +\t\tchar packname[GIT_MAX_HEXSZ + 6];\n> +\t\tconst char *packhash;\n> +\t\tconst int packname_len = the_hash_algo->hexsz + 6;\n\nThe \"+ 6\" here is gross, but not really any more than the bare \"5\" in\nthe original code.\n\nCalling it \"packhash\" made me wonder about this line of code:\n\n> -\t\tif (memcmp(packfile_uris.items[i].string, packname,\n> +\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n\nSurely we need to change more than this if we now have the hash rather\nthan the whole packname? But no, the original code was really just\nstoring the hash in packname.\n\nWhich was rather misleading, but it is not much better after your patch.\nNow packname still just has the packhash, along with the extra keep/pack\nmarker.\n\nWould a more generic name like \"cmd_output\" or something make sense? I\nalso think this would all be much nicer with a strbuf (which would let\nus get rid of the magic numbers), but that is a slightly larger\nrefactor:\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 1e8461d07e..5f94f35c30 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1890,9 +1890,8 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tint created_keep = 0;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 6];\n+\t\tstruct strbuf cmd_output = STRBUF_INIT;\n \t\tconst char *packhash;\n-\t\tconst int packname_len = the_hash_algo->hexsz + 6;\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1910,14 +1909,11 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, packname_len) != packname_len ||\n-\t\t    packname[packname_len - 1] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected pack or keep, TAB, hash, \"\n-\t\t\t    \"then LF in http-fetch output\");\n-\t\tpackname[packname_len - 1] = '\\0';\n-\t\tif (skip_prefix(packname, \"keep\\t\", &packhash))\n+\t\tif (strbuf_read(&cmd_output, cmd.out, 0) < 0)\n+\t\t\tdie(\"failed to read http-fetch output\");\n+\t\tif (skip_prefix(cmd_output.buf, \"keep\\t\", &packhash))\n \t\t\tcreated_keep = 1;\n-\t\telse if (!skip_prefix(packname, \"pack\\t\", &packhash))\n+\t\telse if (!skip_prefix(cmd_output.buf, \"pack\\t\", &packhash))\n \t\t\tdie(\"fetch-pack: expected pack or keep, TAB, hash, \"\n \t\t\t    \"then LF in http-fetch output\");\n \n\n\nBTW, two things that puzzled me while poking at your patch (but neither\nI think are new or the fault of your patch):\n\n  1. git-http-fetch documents --index-pack-args, but the actual option\n     is the singular --index-pack-arg. The caller in fetch-pack\n     obviously uses the one that works.\n\n  2. The code you're touching insists on reading \"keep\" in the output,\n     but wouldn't that depend on feeding \"--keep\" via index_pack_args? I\n     didn't see immediately where we set it, but I certainly don't think\n     your patch could be making anything worse here (since the existing\n     code would have just died upon seeing a \"pack\" line). I think it\n     happens as a side effect in get_pack(), which is...subtle. But\n     again, not anything new.\n\n-Peff\n"},{"id":"548082","messageId":"20260714071331.GA4058163@coredump.intra.peff.net","threadId":"65988","inReplyTo":"20260714071231.GD2516582@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-14T07:13:31Z","receivedAt":"2026-07-14T07:13:33Z","isPatch":true,"body":"On Tue, Jul 14, 2026 at 03:12:31AM -0400, Jeff King wrote:\n\n> Would a more generic name like \"cmd_output\" or something make sense? I\n> also think this would all be much nicer with a strbuf (which would let\n> us get rid of the magic numbers), but that is a slightly larger\n> refactor:\n> \n> diff --git a/fetch-pack.c b/fetch-pack.c\n\nIn case anybody does pursue this, it is obviously missing this bit:\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 5f94f35c30..359740f231 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1935,6 +1935,8 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n \t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n \t\t\t\t\t\t\t packhash));\n+\n+\t\tstrbuf_release(&cmd_output);\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\n\nto avoid a leak.\n\n-Peff\n"},{"id":"548145","messageId":"xmqqcxwptpb0.fsf@gitster.g","threadId":"65988","inReplyTo":"20260714052833.GA2516582@coredump.intra.peff.net","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-14T18:10:43Z","receivedAt":"2026-07-14T18:10:46Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> On Mon, Jul 13, 2026 at 06:58:24PM -0700, Ted Nyman wrote:\n>\n>> > Are there better ways for these processes to coordinate with each\n>> > other? Instead of appending to the file, what if the second process\n>> > uses a predictable temporary name (which we already use) to open a\n>> > new file with O_CREAT | O_EXCL to avoid this redundant work?\n>> \n>> Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL\n>> would prevent concurrent writes, but EEXIST alone would not\n>> distinguish an in-progress download from one left by an earlier\n>> failed or interrupted invocation. The existing .pack.temp name is not\n>> covered by the tmp_* pruning path, so simply waiting for it to\n>> disappear could leave a fetch stuck after a crash.\n>\n> A few thoughts:\n>\n>   - Using O_EXCL makes this essentially a lockfile. So we could apply\n>     the logic used elsewhere for lockfiles, like auto-removing files\n>     with ancient mtimes. Or we could even go all-in with a pid check for\n>     liveness; most of Git's lockfiles don't do that, but at least one\n>     does (the background auto-gc lock).\n>\n>   - If we're not already using a name which is auto-cleaned during\n>     maintenance, we probably ought to be. Leaving aside concurrency\n>     issues, nobody would ever clean up the on-disk cruft.\n>\n>     But of course the original code here is intentionally _not_ using a\n>     name we'd clean up, because it wants to be able to resume an\n>     interrupted transfer.  And you're explicitly breaking that for the\n>     packfile URI case.\n>\n>     Is that a cost we're OK with paying? Fixing it opens up that same\n>     coordination can of worms. You have to tell the difference a\n>     concurrent writer and a previous dead one (whose work you can\n>     resume).\n>\n>     It does feel weird that we'd do one thing for dumb-http and another\n>     for packfile URIs. Wouldn't they suffer from the same concurrency\n>     and resumption problems?\n> ...\n>\n> If we're OK with killing the ability to resume, then yeah, I think it\n> would make sense to start simple and un-break things. And then put a\n> coordination layer on top later (or never if nobody cares enough).\n\nI share that sentiment.  I am not entirely convinced by Ted's\nresponse, since a major goal of the packfile URI feature, as I\nunderstand it, is to allow the use of resumable protocols for\nlarge transfers.  The proposed change deliberately closes the\ndoor on resuming interrupted transfers, whether manually or,\nwith additional code in the future, automatically.\n"},{"id":"548155","messageId":"alaAi4vNwi-KabYV@com-76773","threadId":"65988","inReplyTo":"xmqqcxwptpb0.fsf@gitster.g","subject":"Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-14T18:31:39Z","receivedAt":"2026-07-14T18:31:44Z","isPatch":true,"body":"> I share that sentiment. I am not entirely convinced by Ted's\n> response, since a major goal of the packfile URI feature, as I\n> understand it, is to allow the use of resumable protocols for\n> large transfers.\n\nAgreed. I was too quick to dismiss the loss of resumption.\n\nI'll take another look at preserving the predictable partial pack while\npreventing concurrent writers, including the handoff and stale-file\ncases Peff raised. Dumb HTTP has the same underlying concurrency issue,\nso I'll keep that path in mind as well before sending a reroll.\n\nThanks,\nTed\n"},{"id":"548158","messageId":"alaCQKXKcWr723Ij@com-76773","threadId":"65988","inReplyTo":"20260714071231.GD2516582@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-14T18:38:56Z","receivedAt":"2026-07-14T18:39:08Z","isPatch":true,"body":"> I also think this would all be much nicer with a strbuf (which would\n> let us get rid of the magic numbers), but that is a slightly larger\n> refactor:\n\nUsing a strbuf makes sense. One wrinkle, I think, is that with\ntransfer.fsckobjects enabled, index-pack can emit dangling .gitmodules\nOIDs after the initial pack/keep line, which parse_gitmodules_oids()\nstill needs to read from cmd.out. Would strbuf_getwholeline_fd() be a\nbetter fit here, so we don't consume those with strbuf_read()?\n\nI'll also fix the --index-pack-args documentation while rerolling.\n\nThanks,\nTed\n"},{"id":"548166","messageId":"20260714214709.GA4095533@coredump.intra.peff.net","threadId":"65988","inReplyTo":"alaCQKXKcWr723Ij@com-76773","subject":"Re: [PATCH 2/2] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-14T21:47:09Z","receivedAt":"2026-07-14T21:47:12Z","isPatch":true,"body":"On Tue, Jul 14, 2026 at 11:38:56AM -0700, Ted Nyman wrote:\n\n> > I also think this would all be much nicer with a strbuf (which would\n> > let us get rid of the magic numbers), but that is a slightly larger\n> > refactor:\n> \n> Using a strbuf makes sense. One wrinkle, I think, is that with\n> transfer.fsckobjects enabled, index-pack can emit dangling .gitmodules\n> OIDs after the initial pack/keep line, which parse_gitmodules_oids()\n> still needs to read from cmd.out. Would strbuf_getwholeline_fd() be a\n> better fit here, so we don't consume those with strbuf_read()?\n\nAh, yeah, I didn't think about whether it might have more output. I\n_think_ it actually works just fine with more output because the\nmemcmp() is limited to the hash algo's hex_sz. For the same reason what\nI posted works even though it has the trailing newline.\n\nIt is a bit subtle, though. Using getwholeline_fd would work (though you\nstill have the trailing newline subtlety). Or maybe just using\nstrbuf_setlen() to cut off the output (ironically it is probably more\nefficient to read the whole thing in and then chomp it, since\ngetwholeline_fd will read() one char at a time).\n\nThe \"cleanest\" thing is perhaps xfdopen() followed by strbuf_getline(),\nbut maybe that's overkill.\n\nI'd be happy with any of the solutions. Or even just keeping the magic\nnumbers but maybe with a comment explaining what the heck \"6\" means.\n\n> I'll also fix the --index-pack-args documentation while rerolling.\n\nGreat, thanks.\n\n-Peff\n"},{"id":"548700","messageId":"cover.1784582665.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1783982021.git.tnyman@openai.com","subject":"[PATCH v2 0/2] packfile URIs: support concurrent downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-20T22:33:58Z","receivedAt":"2026-07-20T22:34:03Z","isPatch":true,"body":"Packfile URI and dumb HTTP downloads stage packs at\nobjects/pack/pack-<hash>.pack.temp so an interrupted transfer can\nresume. Two Git processes fetching the same pack into one object\ndatabase can append to that file concurrently, which can corrupt the\ntemporary pack.\n\nFollowing Peff's suggestion in the v1 discussion, the first patch keeps\nthe predictable temporary name but opens the file without append mode.\nEach downloader keeps its own file offset, so overlapping responses for\nthe same pack write the same bytes at the same offsets. This preserves\nresumable downloads and protects both packfile URI and ordinary dumb\nHTTP requests. Concurrent downloaders may still transfer the same\nsuffix; this series avoids adding cross-process coordination and a\nstale-owner policy.\n\nA second downloader can also find that the partial pack has completed\nand request a range starting at EOF. Servers may respond with HTTP 416\nin that case. Treat that response as a completed download and let\nindex-pack validate the pack. Keep the open descriptor for indexing so\nanother downloader can safely remove the temporary path.\n\nThe second patch handles the related .keep race. When another process\nhas already created the keep file, index-pack reports\n\"pack<TAB><hash>\" instead of \"keep<TAB><hash>\". Accept both successful\nforms and remove only keep files created by the current process. Read\nonly the prefix and hash so any following fsck output remains available\nto fetch-pack.\n\nThe tests cover resumption, a completed partial returning 416,\noverlapping 200 and 206 responses, and a pre-existing .keep file.\n\nChanges since v1:\n\n  * Preserve resumability by removing append mode instead of using a\n    unique temporary file for each download.\n  * Handle the EOF-range/416 and concurrent-unlink cases, including the\n    Windows sharing behavior.\n  * Add a deterministic overlapping-download regression test.\n  * Read the pack/keep prefix and hash without consuming later fsck\n    output, as suggested during review.\n  * Correct the stale --index-pack-args documentation and error text;\n    the repeatable --index-pack-arg option is already supported.\n\nThe v1 discussion is at:\n\n  https://lore.kernel.org/git/cover.1783982021.git.tnyman@openai.com/\n\nTed Nyman (2):\n  http: avoid concurrent appends to partial packs\n  fetch-pack: accept \"pack\" output for packfile URIs\n\n Documentation/git-http-fetch.adoc |  13 +-\n fetch-pack.c                      |  33 +++--\n http-fetch.c                      |   7 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  53 ++++---\n t/t5550-http-fetch-dumb.sh        | 223 ++++++++++++++++++++++++++++++\n t/t5702-protocol-v2.sh            |  31 +++++\n 8 files changed, 320 insertions(+), 46 deletions(-)\n\nRange-diff against v1:\n1:  32eb9b0831 ! 1:  160a9b9fd0 http: use unique tempfiles for packfile URI downloads\n    @@ Metadata\n     Author: Ted Nyman <tnyman@openai.com>\n     \n      ## Commit message ##\n    -    http: use unique tempfiles for packfile URI downloads\n    +    http: avoid concurrent appends to partial packs\n     \n    -    Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,\n    -    2020-06-10), packfile URI downloads have been staged at\n    -    objects/pack/pack-<hash>.pack.temp.\n    +    Pack requests stage downloads in a predictable partial-pack file so an\n    +    interrupted transfer can be resumed. Both packfile URI and ordinary dumb\n    +    HTTP requests use this staging path. Opening it in append mode lets\n    +    concurrent fetches interleave their writes, corrupting the pack or\n    +    causing a later fetch to request a range at EOF.\n     \n    -    The path is derived from the advertised pack hash. Two processes\n    -    fetching the same pack into a shared object database therefore open the\n    -    same file for append. Their writes can corrupt the temporary pack. If\n    -    one process arrives after the other has completed the download, it may\n    -    instead try to resume at EOF, which some HTTP servers reject with 416.\n    +    Open the partial pack read-write, seek to its current end, and retain a\n    +    per-descriptor offset for incoming data. Reopen newly created partial\n    +    packs without O_CREAT so Windows permits concurrent unlink, and keep the\n    +    descriptor for index-pack when another downloader removes the staging\n    +    path. Accept HTTP 416 when a partial pack is already complete.\n     \n    -    Use the tempfile API to give direct packfile URI downloads unique\n    -    temporary files. Keep the deterministic path for ordinary dumb HTTP\n    -    pack requests, which use it to resume a partial download left by an\n    -    earlier invocation.\n    -\n    -    This means that a packfile URI download cannot be resumed by a later\n    -    invocation. A retry starts with an empty temporary file instead.\n    -\n    -    Add a test which pauses one process after downloading the pack and\n    -    starts another process using the same object database.\n    +    Exercise resumed transfers, EOF ranges, and overlapping 200 and 206\n    +    responses. Clarify the staging-key documentation and correct the stale\n    +    --index-pack-args spelling in the documentation and error messages; the\n    +    repeatable --index-pack-arg option is already accepted.\n     \n         Signed-off-by: Ted Nyman <tnyman@openai.com>\n    -    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n     \n      ## Documentation/git-http-fetch.adoc ##\n     @@ Documentation/git-http-fetch.adoc: commit-id::\n    @@ Documentation/git-http-fetch.adoc: commit-id::\n     -\tThe hash is used to determine the name of the temporary file and is\n     -\tarbitrary. The output of index-pack is printed to stdout. Requires\n     -\t--index-pack-args.\n    -+\tThe hash is arbitrary. The output of index-pack is printed to stdout.\n    -+\tRequires --index-pack-args.\n    ++\tThe hash is used to determine the name of the temporary file. It need\n    ++\tnot be the pack hash, but it must uniquely identify the pack contents\n    ++\tfor resumption. The output of index-pack is printed to stdout. Requires\n    ++\tone or more --index-pack-arg options.\n    + \n    +---index-pack-args=<args>::\n    +-\tFor internal use only. The command to run on the contents of the\n    +-\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n    ++--index-pack-arg=<arg>::\n    ++\tFor internal use only. An argument to the command run on the contents\n    ++\tof the downloaded pack. This option can be specified multiple times.\n      \n    - --index-pack-args=<args>::\n    - \tFor internal use only. The command to run on the contents of the\n    + --recover::\n    + \tVerify that everything reachable from target is fetched.  Used after\n     \n    - ## http.c ##\n    -@@ http.c: int http_get_info_packs(const char *base_url, struct packfile_list *packs)\n    + ## http-fetch.c ##\n    +@@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,\n      \n    - void release_http_pack_request(struct http_pack_request *preq)\n    - {\n    --\tif (preq->packfile) {\n    -+\tif (preq->tempfile) {\n    -+\t\tdelete_tempfile(&preq->tempfile);\n    -+\t\tpreq->packfile = NULL;\n    -+\t} else if (preq->packfile) {\n    - \t\tfclose(preq->packfile);\n    - \t\tpreq->packfile = NULL;\n    + \tif (start_active_slot(preq->slot)) {\n    + \t\trun_active_slot(preq->slot);\n    +-\t\tif (results.curl_result != CURLE_OK) {\n    ++\t\tif (results.curl_result != CURLE_OK &&\n    ++\t\t    results.http_code != 416) {\n    + \t\t\tstruct url_info url;\n    + \t\t\tchar *nurl = url_normalize(preq->url, &url);\n    + \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\n    +@@ http-fetch.c: int cmd_main(int argc, const char **argv)\n    + \n    + \tif (packfile) {\n    + \t\tif (!index_pack_args.nr)\n    +-\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n    ++\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n    + \n    + \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n    + \t\t\t\t      index_pack_args.v);\n    +@@ http-fetch.c: int cmd_main(int argc, const char **argv)\n      \t}\n    + \n    + \tif (index_pack_args.nr)\n    +-\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n    ++\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n    + \n    + \tif (commits_on_stdin) {\n    + \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n    +\n    + ## http-push.c ##\n    +@@ http-push.c: static void finish_request(struct transfer_request *request)\n    + \n    + \t} else if (request->state == RUN_FETCH_PACKED) {\n    + \t\tint fail = 1;\n    +-\t\tif (request->curl_result != CURLE_OK) {\n    ++\t\tif (request->curl_result != CURLE_OK &&\n    ++\t\t    request->http_code != 416) {\n    + \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n    + \t\t\t\trequest->url, curl_errorstr);\n    + \t\t} else {\n    +\n    + ## http-walker.c ##\n    +@@ http-walker.c: static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n    + \n    + \tif (start_active_slot(preq->slot)) {\n    + \t\trun_active_slot(preq->slot);\n    +-\t\tif (results.curl_result != CURLE_OK) {\n    ++\t\tif (results.curl_result != CURLE_OK &&\n    ++\t\t    results.http_code != 416) {\n    + \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n    + \t\t\t      curl_errorstr);\n    + \t\t\tgoto abort;\n    +\n    + ## http.c ##\n     @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)\n      \tint tmpfile_fd;\n      \tint ret = 0;\n      \n    --\tfclose(preq->packfile);\n    -+\tif (preq->tempfile)\n    -+\t\tclose_tempfile_gently(preq->tempfile);\n    -+\telse\n    -+\t\tfclose(preq->packfile);\n    ++\t/* Another downloader may unlink the staging path while we index it. */\n    ++\ttmpfile_fd = xdup(fileno(preq->packfile));\n    + \tfclose(preq->packfile);\n      \tpreq->packfile = NULL;\n    +-\n    +-\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n    ++\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n    ++\t\tdie_errno(\"unable to seek local file %s for pack\",\n    ++\t\t\t  preq->tmpfile.buf);\n      \n    - \ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n    + \tip.git_cmd = 1;\n    + \tip.in = tmpfile_fd;\n     @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)\n    + \telse\n    + \t\tip.no_stdout = 1;\n      \n    - cleanup:\n    - \tclose(tmpfile_fd);\n    --\tunlink(preq->tmpfile.buf);\n    -+\tif (preq->tempfile)\n    -+\t\tdelete_tempfile(&preq->tempfile);\n    -+\telse\n    -+\t\tunlink(preq->tmpfile.buf);\n    +-\tif (run_command(&ip)) {\n    ++\tif (run_command(&ip))\n    + \t\tret = -1;\n    +-\t\tgoto cleanup;\n    +-\t}\n    +-\n    +-cleanup:\n    +-\tclose(tmpfile_fd);\n    + \tunlink(preq->tmpfile.buf);\n      \treturn ret;\n      }\n    - \n    -@@ http.c: void http_install_packfile(struct packed_git *p,\n    - \tpackfile_store_add_pack(files->packed, p);\n    - }\n    - \n    --struct http_pack_request *new_http_pack_request(\n    --\tconst unsigned char *packed_git_hash, const char *base_url) {\n    --\n    --\tstruct strbuf buf = STRBUF_INIT;\n    --\n    --\tend_url_with_slash(&buf, base_url);\n    --\tstrbuf_addf(&buf, \"objects/pack/pack-%s.pack\",\n    --\t\thash_to_hex(packed_git_hash));\n    --\treturn new_direct_http_pack_request(packed_git_hash,\n    --\t\t\t\t\t    strbuf_detach(&buf, NULL));\n    --}\n    --\n    --struct http_pack_request *new_direct_http_pack_request(\n    --\tconst unsigned char *packed_git_hash, char *url)\n    -+static struct http_pack_request *new_http_pack_request_for_url(\n    -+\tconst unsigned char *packed_git_hash, char *url, int resumable)\n    +@@ http.c: struct http_pack_request *new_http_pack_request(\n    + struct http_pack_request *new_direct_http_pack_request(\n    + \tconst unsigned char *packed_git_hash, char *url)\n      {\n    - \toff_t prev_posn = 0;\n    +-\toff_t prev_posn = 0;\n    ++\toff_t prev_posn;\n      \tstruct http_pack_request *preq;\n    -@@ http.c: struct http_pack_request *new_direct_http_pack_request(\n    ++\tint fd;\n      \n    + \tCALLOC_ARRAY(preq, 1);\n    + \tstrbuf_init(&preq->tmpfile, 0);\n    +-\n      \tpreq->url = url;\n      \n    --\todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n    --\tstrbuf_addstr(&preq->tmpfile, \".temp\");\n    + \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n    + \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n     -\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n    -+\tif (resumable) {\n    -+\t\todb_pack_name(the_repository, &preq->tmpfile,\n    -+\t\t\t      packed_git_hash, \"pack\");\n    -+\t\tstrbuf_addstr(&preq->tmpfile, \".temp\");\n    -+\t\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n    -+\t} else {\n    -+\t\tstrbuf_addf(&preq->tmpfile, \"%s/pack/tmp_pack_XXXXXX\",\n    -+\t\t\t    repo_get_object_directory(the_repository));\n    -+\t\tpreq->tempfile = mks_tempfile_m(preq->tmpfile.buf, 0444);\n    -+\t\tif (preq->tempfile) {\n    -+\t\t\tstrbuf_reset(&preq->tmpfile);\n    -+\t\t\tstrbuf_addstr(&preq->tmpfile,\n    -+\t\t\t\t      get_tempfile_path(preq->tempfile));\n    -+\t\t\tpreq->packfile = fdopen_tempfile(preq->tempfile, \"w\");\n    +-\tif (!preq->packfile) {\n    +-\t\terror(\"Unable to open local file %s for pack\",\n    +-\t\t      preq->tmpfile.buf);\n    ++\t/* Reopen without O_CREAT so MinGW permits another writer to unlink it. */\n    ++\tfor (;;) {\n    ++\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n    ++\t\tif (fd >= 0 || errno != ENOENT)\n    ++\t\t\tbreak;\n    ++\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n    ++\t\tif (fd >= 0) {\n    ++\t\t\tclose(fd);\n    ++\t\t\tcontinue;\n     +\t\t}\n    ++\t\tif (errno != EEXIST)\n    ++\t\t\tbreak;\n    ++\t}\n    ++\tif (fd < 0) {\n    ++\t\terror_errno(\"unable to open local file %s for pack\",\n    ++\t\t\t    preq->tmpfile.buf);\n    ++\t\tgoto abort;\n     +\t}\n    - \tif (!preq->packfile) {\n    - \t\terror(\"Unable to open local file %s for pack\",\n    - \t\t      preq->tmpfile.buf);\n    ++\tprev_posn = lseek(fd, 0, SEEK_END);\n    ++\tif (prev_posn < 0) {\n    ++\t\terror_errno(\"unable to seek local file %s for pack\",\n    ++\t\t\t    preq->tmpfile.buf);\n    ++\t\tclose(fd);\n    + \t\tgoto abort;\n    + \t}\n    ++\tpreq->packfile = xfdopen(fd, \"w\");\n    + \n    + \tpreq->slot = get_active_slot();\n    + \tpreq->headers = object_request_headers();\n     @@ http.c: struct http_pack_request *new_direct_http_pack_request(\n    - \t * If there is data present from a previous transfer attempt,\n    - \t * resume where it left off\n    - \t */\n    + \tcurl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);\n    + \tcurl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);\n    + \n    +-\t/*\n    +-\t * If there is data present from a previous transfer attempt,\n    +-\t * resume where it left off\n    +-\t */\n     -\tprev_posn = ftello(preq->packfile);\n     -\tif (prev_posn>0) {\n    -+\tif (resumable)\n    -+\t\tprev_posn = ftello(preq->packfile);\n     +\tif (prev_posn > 0) {\n      \t\tif (http_is_verbose)\n      \t\t\tfprintf(stderr,\n      \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\n    -@@ http.c: struct http_pack_request *new_direct_http_pack_request(\n    - \treturn preq;\n    - \n    - abort:\n    --\tstrbuf_release(&preq->tmpfile);\n    --\tfree(preq->url);\n    --\tfree(preq);\n    -+\trelease_http_pack_request(preq);\n    - \treturn NULL;\n    - }\n    - \n    -+struct http_pack_request *new_http_pack_request(\n    -+\tconst unsigned char *packed_git_hash, const char *base_url)\n    -+{\n    -+\tstruct strbuf buf = STRBUF_INIT;\n    -+\n    -+\tend_url_with_slash(&buf, base_url);\n    -+\tstrbuf_addf(&buf, \"objects/pack/pack-%s.pack\",\n    -+\t\thash_to_hex(packed_git_hash));\n    -+\treturn new_http_pack_request_for_url(packed_git_hash,\n    -+\t\t\t\t\t     strbuf_detach(&buf, NULL), 1);\n    -+}\n    -+\n    -+struct http_pack_request *new_direct_http_pack_request(\n    -+\tconst unsigned char *packed_git_hash, char *url)\n    -+{\n    -+\treturn new_http_pack_request_for_url(packed_git_hash, url, 0);\n    -+}\n    -+\n    - /* Helpers for fetching objects (loose) */\n    - static size_t fwrite_sha1_file(char *ptr, size_t eltsize, size_t nmemb,\n    - \t\t\t       void *data)\n    -\n    - ## http.h ##\n    -@@ http.h: struct http_pack_request {\n    - \n    - \tFILE *packfile;\n    - \tstruct strbuf tmpfile;\n    -+\tstruct tempfile *tempfile;\n    - \tstruct active_request_slot *slot;\n    - \tstruct curl_slist *headers;\n    - };\n     \n      ## t/t5550-http-fetch-dumb.sh ##\n     @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n      \tgit -C packfileclient cat-file -e \"$HASH\"\n      '\n      \n    -+test_expect_success PIPE 'concurrent http-fetch --packfile' '\n    ++test_expect_success 'http-fetch --packfile resumes a partial download' '\n    ++\tgit init packfileclient-resume &&\n    ++\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n    ++\t\tls objects/pack/pack-*.pack) &&\n    ++\ttmpfile=\"packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n    ++\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n    ++\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/resume.trace\" \\\n    ++\tgit -C packfileclient-resume http-fetch --packfile=\"$ARBITRARY\" \\\n    ++\t\t--index-pack-arg=index-pack --index-pack-arg=--stdin \\\n    ++\t\t--index-pack-arg=--keep \\\n    ++\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n    ++\ttest_grep \"Range: bytes=64-\" resume.trace &&\n    ++\ttest_path_is_missing \"$tmpfile\" &&\n    ++\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n    ++'\n    ++\n    ++test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n     +\tgit init packfileclient-concurrent &&\n    -+\tHASH=$(git -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git rev-parse HEAD) &&\n     +\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n     +\t\tls objects/pack/pack-*.pack) &&\n     +\tpackhash=$(basename \"$p\" .pack) &&\n     +\tpackhash=${packhash#pack-} &&\n    -+\n    ++\ttmpfile=\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n    ++\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n     +\tmkfifo first-ready first-continue &&\n     +\texec 8<>first-ready &&\n     +\texec 9<>first-continue &&\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n     +\texec git index-pack \"$@\"\n     +\tEOF\n    -+\n    -+\t# Hold the first download before it is indexed, so that the second\n    -+\t# download installs the pack first.\n     +\t{\n     +\t\t(\n     +\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n     +\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n     +\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n    -+\t\t\tgit -C packfileclient-concurrent http-fetch \\\n    -+\t\t\t\t--packfile=\"$packhash\" \\\n    ++\t\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/first.trace\" \\\n    ++\t\t\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n     +\t\t\t\t--index-pack-arg=wait-index-pack \\\n    -+\t\t\t\t--index-pack-arg=--stdin \\\n    -+\t\t\t\t--index-pack-arg=--keep \\\n    ++\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n     +\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n     +\t\t\tthen\n     +\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\t} &&\n     +\ttest_when_finished \"\n     +\t\techo continue >&9\n    ++\t\tkill $first_pid 2>/dev/null || :\n     +\t\twait $first_pid 2>/dev/null || :\n     +\t\texec 8>&-\n     +\t\texec 9>&-\n     +\t\trm -f first-ready first-continue git-wait-index-pack\n     +\t\" &&\n    -+\n     +\tread ready <&8 &&\n     +\ttest \"$ready\" = ready &&\n    -+\tgit -C packfileclient-concurrent http-fetch \\\n    -+\t\t--packfile=\"$packhash\" \\\n    ++\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/second.trace\" \\\n    ++\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n     +\t\t--index-pack-arg=index-pack \\\n    -+\t\t--index-pack-arg=--stdin \\\n    -+\t\t--index-pack-arg=--keep \\\n    ++\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n     +\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n     +\techo continue >&9 &&\n     +\twait \"$first_pid\" &&\n    -+\n     +\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n     +\ttest_cmp expect first.out &&\n     +\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n     +\ttest_cmp expect second.out &&\n    -+\ttest_path_is_missing \\\n    -+\t\t\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n    -+\tfind packfileclient-concurrent/.git/objects/pack \\\n    -+\t\t-name \"tmp_pack_*\" -print >tmpfiles &&\n    -+\ttest_must_be_empty tmpfiles &&\n    ++\ttest_grep \"Range: bytes=64-\" first.trace &&\n    ++\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n    ++\ttest_grep \"HTTP/[0-9.]* 416\" second.trace &&\n    ++\ttest_path_is_missing \"$tmpfile\" &&\n     +\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n     +'\n    ++\n    ++test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n    ++\tgit init packfileclient-overlap &&\n    ++\tblob=$(test-tool genrandom pack-overlap 2m |\n    ++\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n    ++\t\t\thash-object -w --stdin) &&\n    ++\tpackhash=$(printf \"%s\\n\" \"$blob\" |\n    ++\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n    ++\t\t\tpack-objects \"$TRASH_DIRECTORY/overlap-pack\") &&\n    ++\tpack=\"$TRASH_DIRECTORY/overlap-pack-$packhash.pack\" &&\n    ++\ttmpfile=\"packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp\" &&\n    ++\tmkfifo server-ready first-ready &&\n    ++\texec 7<>server-ready &&\n    ++\texec 8<>first-ready &&\n    ++\twrite_script slow-pack-server \"$PERL_PATH\" <<-\\EOF &&\n    ++\tuse strict;\n    ++\tuse warnings;\n    ++\tuse IO::Socket::INET;\n    ++\n    ++\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n    ++\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n    ++\tmy $pack = do { local $/; <$in> };\n    ++\tclose($in) or die \"close $packfile: $!\";\n    ++\tmy $server = IO::Socket::INET->new(LocalAddr => \"127.0.0.1\",\n    ++\t\tLocalPort => 0, Proto => \"tcp\", Listen => 2, ReuseAddr => 1)\n    ++\t\tor die \"listen: $!\";\n    ++\n    ++\tsub signal_ready {\n    ++\t\tmy ($file, $value) = @_;\n    ++\t\topen(my $out, \">\", $file) or die \"open $file: $!\";\n    ++\t\tprint $out \"$value\\n\" or die \"write $file: $!\";\n    ++\t\tclose($out) or die \"close $file: $!\";\n    ++\t}\n    ++\n    ++\tsub write_all {\n    ++\t\tmy ($out, $data) = @_;\n    ++\t\tmy $offset = 0;\n    ++\t\twhile ($offset < length($data)) {\n    ++\t\t\tmy $written = syswrite($out, $data,\n    ++\t\t\t\tlength($data) - $offset, $offset);\n    ++\t\t\tdefined($written) && $written or die \"write response: $!\";\n    ++\t\t\t$offset += $written;\n    ++\t\t}\n    ++\t}\n    ++\n    ++\tsub start_response {\n    ++\t\tmy $out = $server->accept() or die \"accept: $!\";\n    ++\t\t<$out> or die \"read request: $!\";\n    ++\t\tmy $start = 0;\n    ++\t\twhile (<$out>) {\n    ++\t\t\tlast if /^\\r?\\n$/;\n    ++\t\t\t$start = $1 if /^Range: bytes=(\\d+)-/i;\n    ++\t\t}\n    ++\t\t$start < length($pack) or die \"invalid range $start\";\n    ++\t\tmy $length = length($pack) - $start;\n    ++\t\tmy $middle = int($length / 2);\n    ++\t\tmy $status = $start ? \"206 Partial Content\" : \"200 OK\";\n    ++\t\tmy $headers = \"HTTP/1.1 $status\\r\\n\" .\n    ++\t\t\t\"Content-Length: $length\\r\\n\" .\n    ++\t\t\t($start ? \"Content-Range: bytes $start-\" .\n    ++\t\t\t\t(length($pack) - 1) . \"/\" . length($pack) . \"\\r\\n\" : \"\") .\n    ++\t\t\t\"Connection: close\\r\\n\\r\\n\";\n    ++\t\twrite_all($out, $headers);\n    ++\t\twrite_all($out, substr($pack, $start, $middle));\n    ++\t\treturn ($out, $start + $middle);\n    ++\t}\n    ++\n    ++\tsignal_ready($server_ready, $server->sockport());\n    ++\tmy ($first, $first_pos) = start_response();\n    ++\tsignal_ready($first_ready, \"ready\");\n    ++\tmy ($second, $second_pos) = start_response();\n    ++\twrite_all($first, substr($pack, $first_pos));\n    ++\twrite_all($second, substr($pack, $second_pos));\n    ++\tclose($first) or die \"close first response: $!\";\n    ++\tclose($second) or die \"close second response: $!\";\n    ++\tEOF\n    ++\t{\n    ++\t\t(\n    ++\t\t\tif ! \"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n    ++\t\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n    ++\t\t\t\t\"$TRASH_DIRECTORY/first-ready\"\n    ++\t\t\tthen\n    ++\t\t\t\techo failed >\"$TRASH_DIRECTORY/server-ready\" &&\n    ++\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n    ++\t\t\t\texit 1\n    ++\t\t\tfi\n    ++\t\t) >server.log 2>&1 &\n    ++\t\tserver_pid=$!\n    ++\t} &&\n    ++\ttest_when_finished \"\n    ++\t\tkill $server_pid 2>/dev/null || :\n    ++\t\twait $server_pid 2>/dev/null || :\n    ++\t\texec 7>&-\n    ++\t\texec 8>&-\n    ++\t\trm -f server-ready first-ready slow-pack-server\n    ++\t\" &&\n    ++\tread port <&7 &&\n    ++\turl=\"http://127.0.0.1:$port/pack\" &&\n    ++\t{\n    ++\t\t(\n    ++\t\t\tif ! GIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-first.trace\" \\\n    ++\t\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n    ++\t\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n    ++\t\t\t\t--index-pack-arg=index-pack \\\n    ++\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n    ++\t\t\t\t\"$url\" >first.out\n    ++\t\t\tthen\n    ++\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n    ++\t\t\t\texit 1\n    ++\t\t\tfi\n    ++\t\t) &\n    ++\t\tfirst_pid=$!\n    ++\t} &&\n    ++\ttest_when_finished \"\n    ++\t\tkill $first_pid 2>/dev/null || :\n    ++\t\twait $first_pid 2>/dev/null || :\n    ++\t\" &&\n    ++\tread ready <&8 &&\n    ++\ttest \"$ready\" = ready &&\n    ++\ttest_path_is_file \"$tmpfile\" &&\n    ++\ttest -s \"$tmpfile\" &&\n    ++\t{\n    ++\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n    ++\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n    ++\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n    ++\t\t\t--index-pack-arg=index-pack \\\n    ++\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n    ++\t\t\t\"$url\" >second.out &\n    ++\t\tsecond_pid=$!\n    ++\t} &&\n    ++\ttest_when_finished \"\n    ++\t\tkill $second_pid 2>/dev/null || :\n    ++\t\twait $second_pid 2>/dev/null || :\n    ++\t\" &&\n    ++\twait \"$server_pid\" &&\n    ++\twait \"$first_pid\" &&\n    ++\twait \"$second_pid\" &&\n    ++\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n    ++\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n    ++\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n    ++\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n    ++\tsort first.out second.out >actual &&\n    ++\ttest_cmp expect actual &&\n    ++\ttest_path_is_missing \"$tmpfile\" &&\n    ++\tgit -C packfileclient-overlap cat-file -e \"$blob\"\n    ++'\n     +\n      test_expect_success 'fetch notices corrupt pack' '\n      \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n      \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n    -@@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile with corrupt pack' '\n    - \tgit init packfileclient &&\n    - \tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git && ls objects/pack/pack-*.pack) &&\n    - \ttest_must_fail git -C packfileclient http-fetch --packfile \\\n    --\t\t\"$HTTPD_URL\"/dumb/repo_bad1.git/$p\n    -+\t\t\"$HTTPD_URL\"/dumb/repo_bad1.git/$p &&\n    -+\tfind packfileclient/.git/objects/pack -name \"tmp_pack_*\" -print >tmpfiles &&\n    -+\ttest_must_be_empty tmpfiles\n    - '\n    - \n    - test_expect_success 'fetch notices corrupt idx' '\n2:  e73de423f0 ! 2:  9b41d4ddb3 fetch-pack: accept \"pack\" output for packfile URIs\n    @@ Metadata\n      ## Commit message ##\n         fetch-pack: accept \"pack\" output for packfile URIs\n     \n    -    When \"index-pack --keep\" creates a .keep file, it reports\n    -    \"keep<TAB><hash>\". If the file already exists, index-pack leaves it\n    -    untouched and reports \"pack<TAB><hash>\" instead.\n    +    When index-pack finds an existing keep file it reports pack rather than\n    +    keep. Accept either result from http-fetch, and only register a keep\n    +    lockfile when this fetch created it.\n     \n    -    Since dd4b732df7 (upload-pack: send part of packfile response as uri,\n    -    2020-06-10), fetch-pack has accepted only the \"keep\" form for packs\n    -    downloaded through packfile URIs. A concurrent fetch can install the\n    -    same pack and create its .keep file before another process reaches\n    -    index-pack. The latter process then fails even though index-pack\n    -    completed successfully.\n    -\n    -    Accept both successful forms. Add a path to pack_lockfiles only for the\n    -    \"keep\" form, so cleanup removes only a keep file created by the current\n    -    process and preserves a pre-existing one.\n    -\n    -    Add a regression test which pre-creates a keep file and verifies that a\n    -    fetch succeeds without changing it.\n    +    Read the pack/keep prefix and hash without consuming any following fsck\n    +    output, validate the reported pack hash against the advertised hash, and\n    +    exercise a packfile URI fetch with a pre-existing keep file.\n     \n         Signed-off-by: Ted Nyman <tnyman@openai.com>\n    -    Signed-off-by: Junio C Hamano <gitster@pobox.com>\n     \n      ## fetch-pack.c ##\n     @@ fetch-pack.c: static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n      \t}\n      \n      \tfor (i = 0; i < packfile_uris.nr; i++) {\n    -+\t\tint created_keep = 0;\n    ++\t\tbool created_keep;\n      \t\tint j;\n      \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n     -\t\tchar packname[GIT_MAX_HEXSZ + 1];\n    -+\t\tchar packname[GIT_MAX_HEXSZ + 6];\n    -+\t\tconst char *packhash;\n    -+\t\tconst int packname_len = the_hash_algo->hexsz + 6;\n    ++\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n      \t\tconst char *uri = packfile_uris.items[i].string +\n      \t\t\tthe_hash_algo->hexsz + 1;\n      \n    @@ fetch-pack.c: static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n     -\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n     -\t\t    memcmp(packname, \"keep\\t\", 5))\n     -\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n    --\n    ++\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n    ++\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n    ++\t\t     memcmp(packhash, \"pack\\t\", 5)))\n    ++\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n    ++\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n    + \n     -\t\tif (read_in_full(cmd.out, packname,\n     -\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n     -\t\t    packname[the_hash_algo->hexsz] != '\\n')\n     -\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n     -\n     -\t\tpackname[the_hash_algo->hexsz] = '\\0';\n    -+\t\tif (read_in_full(cmd.out, packname, packname_len) != packname_len ||\n    -+\t\t    packname[packname_len - 1] != '\\n')\n    -+\t\t\tdie(\"fetch-pack: expected pack or keep, TAB, hash, \"\n    -+\t\t\t    \"then LF in http-fetch output\");\n    -+\t\tpackname[packname_len - 1] = '\\0';\n    -+\t\tif (skip_prefix(packname, \"keep\\t\", &packhash))\n    -+\t\t\tcreated_keep = 1;\n    -+\t\telse if (!skip_prefix(packname, \"pack\\t\", &packhash))\n    -+\t\t\tdie(\"fetch-pack: expected pack or keep, TAB, hash, \"\n    -+\t\t\t    \"then LF in http-fetch output\");\n    ++\t\tif (read_in_full(cmd.out, packhash,\n    ++\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n    ++\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n    ++\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n    ++\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n      \n      \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n      \n\nbase-commit: f60db8d575adb79761d363e026fb49bddf330c73\n-- \n2.55.0.125.g9b41d4ddb3\n"},{"id":"548701","messageId":"160a9b9fd0982dadfbf6f8fbb378d1a3e9173698.1784582665.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784582665.git.tnyman@openai.com","subject":"[PATCH v2 1/2] http: avoid concurrent appends to partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-20T22:33:59Z","receivedAt":"2026-07-20T22:34:04Z","isPatch":true,"body":"Pack requests stage downloads in a predictable partial-pack file so an\ninterrupted transfer can be resumed. Both packfile URI and ordinary dumb\nHTTP requests use this staging path. Opening it in append mode lets\nconcurrent fetches interleave their writes, corrupting the pack or\ncausing a later fetch to request a range at EOF.\n\nOpen the partial pack read-write, seek to its current end, and retain a\nper-descriptor offset for incoming data. Reopen newly created partial\npacks without O_CREAT so Windows permits concurrent unlink, and keep the\ndescriptor for index-pack when another downloader removes the staging\npath. Accept HTTP 416 when a partial pack is already complete.\n\nExercise resumed transfers, EOF ranges, and overlapping 200 and 206\nresponses. Clarify the staging-key documentation and correct the stale\n--index-pack-args spelling in the documentation and error messages; the\nrepeatable --index-pack-arg option is already accepted.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc |  13 +-\n http-fetch.c                      |   7 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  53 ++++---\n t/t5550-http-fetch-dumb.sh        | 223 ++++++++++++++++++++++++++++++\n 6 files changed, 271 insertions(+), 31 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 2200f073c4..60ca91cf3a 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -48,13 +48,14 @@ commit-id::\n \tline (which is not expected in\n \tthis case), 'git http-fetch' fetches the packfile directly at the given\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n-\tThe hash is used to determine the name of the temporary file and is\n-\tarbitrary. The output of index-pack is printed to stdout. Requires\n-\t--index-pack-args.\n+\tThe hash is used to determine the name of the temporary file. It need\n+\tnot be the pack hash, but it must uniquely identify the pack contents\n+\tfor resumption. The output of index-pack is printed to stdout. Requires\n+\tone or more --index-pack-arg options.\n \n---index-pack-args=<args>::\n-\tFor internal use only. The command to run on the contents of the\n-\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n+--index-pack-arg=<arg>::\n+\tFor internal use only. An argument to the command run on the contents\n+\tof the downloaded pack. This option can be specified multiple times.\n \n --recover::\n \tVerify that everything reachable from target is fetched.  Used after\ndiff --git a/http-fetch.c b/http-fetch.c\nindex f9b6ecb061..05f68f306a 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\tstruct url_info url;\n \t\t\tchar *nurl = url_normalize(preq->url, &url);\n \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\n@@ -155,7 +156,7 @@ int cmd_main(int argc, const char **argv)\n \n \tif (packfile) {\n \t\tif (!index_pack_args.nr)\n-\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n+\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n \n \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n \t\t\t\t      index_pack_args.v);\n@@ -164,7 +165,7 @@ int cmd_main(int argc, const char **argv)\n \t}\n \n \tif (index_pack_args.nr)\n-\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n+\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n \n \tif (commits_on_stdin) {\n \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\ndiff --git a/http-push.c b/http-push.c\nindex 3c23cbba27..03dc8102a1 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)\n \n \t} else if (request->state == RUN_FETCH_PACKED) {\n \t\tint fail = 1;\n-\t\tif (request->curl_result != CURLE_OK) {\n+\t\tif (request->curl_result != CURLE_OK &&\n+\t\t    request->http_code != 416) {\n \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n \t\t\t\trequest->url, curl_errorstr);\n \t\t} else {\ndiff --git a/http-walker.c b/http-walker.c\nindex b58a3b2a92..abafca84d6 100644\n--- a/http-walker.c\n+++ b/http-walker.c\n@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n \t\t\t      curl_errorstr);\n \t\t\tgoto abort;\ndiff --git a/http.c b/http.c\nindex b4e7b8d00b..9b9f4efe28 100644\n--- a/http.c\n+++ b/http.c\n@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \tint tmpfile_fd;\n \tint ret = 0;\n \n+\t/* Another downloader may unlink the staging path while we index it. */\n+\ttmpfile_fd = xdup(fileno(preq->packfile));\n \tfclose(preq->packfile);\n \tpreq->packfile = NULL;\n-\n-\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n+\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n+\t\tdie_errno(\"unable to seek local file %s for pack\",\n+\t\t\t  preq->tmpfile.buf);\n \n \tip.git_cmd = 1;\n \tip.in = tmpfile_fd;\n@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \telse\n \t\tip.no_stdout = 1;\n \n-\tif (run_command(&ip)) {\n+\tif (run_command(&ip))\n \t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\n-cleanup:\n-\tclose(tmpfile_fd);\n \tunlink(preq->tmpfile.buf);\n \treturn ret;\n }\n@@ -2738,22 +2736,42 @@ struct http_pack_request *new_http_pack_request(\n struct http_pack_request *new_direct_http_pack_request(\n \tconst unsigned char *packed_git_hash, char *url)\n {\n-\toff_t prev_posn = 0;\n+\toff_t prev_posn;\n \tstruct http_pack_request *preq;\n+\tint fd;\n \n \tCALLOC_ARRAY(preq, 1);\n \tstrbuf_init(&preq->tmpfile, 0);\n-\n \tpreq->url = url;\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n-\tif (!preq->packfile) {\n-\t\terror(\"Unable to open local file %s for pack\",\n-\t\t      preq->tmpfile.buf);\n+\t/* Reopen without O_CREAT so MinGW permits another writer to unlink it. */\n+\tfor (;;) {\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n+\t\tif (fd >= 0 || errno != ENOENT)\n+\t\t\tbreak;\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n+\t\tif (fd >= 0) {\n+\t\t\tclose(fd);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (errno != EEXIST)\n+\t\t\tbreak;\n+\t}\n+\tif (fd < 0) {\n+\t\terror_errno(\"unable to open local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tgoto abort;\n+\t}\n+\tprev_posn = lseek(fd, 0, SEEK_END);\n+\tif (prev_posn < 0) {\n+\t\terror_errno(\"unable to seek local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tclose(fd);\n \t\tgoto abort;\n \t}\n+\tpreq->packfile = xfdopen(fd, \"w\");\n \n \tpreq->slot = get_active_slot();\n \tpreq->headers = object_request_headers();\n@@ -2762,12 +2780,7 @@ struct http_pack_request *new_direct_http_pack_request(\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);\n \n-\t/*\n-\t * If there is data present from a previous transfer attempt,\n-\t * resume where it left off\n-\t */\n-\tprev_posn = ftello(preq->packfile);\n-\tif (prev_posn>0) {\n+\tif (prev_posn > 0) {\n \t\tif (http_is_verbose)\n \t\t\tfprintf(stderr,\n \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex b0080bf204..7acae96a96 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -293,6 +293,229 @@ test_expect_success 'http-fetch --packfile' '\n \tgit -C packfileclient cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile resumes a partial download' '\n+\tgit init packfileclient-resume &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/resume.trace\" \\\n+\tgit -C packfileclient-resume http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=index-pack --index-pack-arg=--stdin \\\n+\t\t--index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_grep \"Range: bytes=64-\" resume.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n+\tgit init packfileclient-concurrent &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\tpackhash=$(basename \"$p\" .pack) &&\n+\tpackhash=${packhash#pack-} &&\n+\ttmpfile=\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tmkfifo first-ready first-continue &&\n+\texec 8<>first-ready &&\n+\texec 9<>first-continue &&\n+\twrite_script git-wait-index-pack <<-\\EOF &&\n+\techo ready >\"$GIT_TEST_WAIT_READY\" &&\n+\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n+\texec git index-pack \"$@\"\n+\tEOF\n+\t{\n+\t\t(\n+\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n+\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n+\t\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/first.trace\" \\\n+\t\t\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=wait-index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\techo continue >&9\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\texec 8>&-\n+\t\texec 9>&-\n+\t\trm -f first-ready first-continue git-wait-index-pack\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/second.trace\" \\\n+\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t--index-pack-arg=index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n+\techo continue >&9 &&\n+\twait \"$first_pid\" &&\n+\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect first.out &&\n+\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect second.out &&\n+\ttest_grep \"Range: bytes=64-\" first.trace &&\n+\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n+\ttest_grep \"HTTP/[0-9.]* 416\" second.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n+\tgit init packfileclient-overlap &&\n+\tblob=$(test-tool genrandom pack-overlap 2m |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\thash-object -w --stdin) &&\n+\tpackhash=$(printf \"%s\\n\" \"$blob\" |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\tpack-objects \"$TRASH_DIRECTORY/overlap-pack\") &&\n+\tpack=\"$TRASH_DIRECTORY/overlap-pack-$packhash.pack\" &&\n+\ttmpfile=\"packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tmkfifo server-ready first-ready &&\n+\texec 7<>server-ready &&\n+\texec 8<>first-ready &&\n+\twrite_script slow-pack-server \"$PERL_PATH\" <<-\\EOF &&\n+\tuse strict;\n+\tuse warnings;\n+\tuse IO::Socket::INET;\n+\n+\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n+\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n+\tmy $pack = do { local $/; <$in> };\n+\tclose($in) or die \"close $packfile: $!\";\n+\tmy $server = IO::Socket::INET->new(LocalAddr => \"127.0.0.1\",\n+\t\tLocalPort => 0, Proto => \"tcp\", Listen => 2, ReuseAddr => 1)\n+\t\tor die \"listen: $!\";\n+\n+\tsub signal_ready {\n+\t\tmy ($file, $value) = @_;\n+\t\topen(my $out, \">\", $file) or die \"open $file: $!\";\n+\t\tprint $out \"$value\\n\" or die \"write $file: $!\";\n+\t\tclose($out) or die \"close $file: $!\";\n+\t}\n+\n+\tsub write_all {\n+\t\tmy ($out, $data) = @_;\n+\t\tmy $offset = 0;\n+\t\twhile ($offset < length($data)) {\n+\t\t\tmy $written = syswrite($out, $data,\n+\t\t\t\tlength($data) - $offset, $offset);\n+\t\t\tdefined($written) && $written or die \"write response: $!\";\n+\t\t\t$offset += $written;\n+\t\t}\n+\t}\n+\n+\tsub start_response {\n+\t\tmy $out = $server->accept() or die \"accept: $!\";\n+\t\t<$out> or die \"read request: $!\";\n+\t\tmy $start = 0;\n+\t\twhile (<$out>) {\n+\t\t\tlast if /^\\r?\\n$/;\n+\t\t\t$start = $1 if /^Range: bytes=(\\d+)-/i;\n+\t\t}\n+\t\t$start < length($pack) or die \"invalid range $start\";\n+\t\tmy $length = length($pack) - $start;\n+\t\tmy $middle = int($length / 2);\n+\t\tmy $status = $start ? \"206 Partial Content\" : \"200 OK\";\n+\t\tmy $headers = \"HTTP/1.1 $status\\r\\n\" .\n+\t\t\t\"Content-Length: $length\\r\\n\" .\n+\t\t\t($start ? \"Content-Range: bytes $start-\" .\n+\t\t\t\t(length($pack) - 1) . \"/\" . length($pack) . \"\\r\\n\" : \"\") .\n+\t\t\t\"Connection: close\\r\\n\\r\\n\";\n+\t\twrite_all($out, $headers);\n+\t\twrite_all($out, substr($pack, $start, $middle));\n+\t\treturn ($out, $start + $middle);\n+\t}\n+\n+\tsignal_ready($server_ready, $server->sockport());\n+\tmy ($first, $first_pos) = start_response();\n+\tsignal_ready($first_ready, \"ready\");\n+\tmy ($second, $second_pos) = start_response();\n+\twrite_all($first, substr($pack, $first_pos));\n+\twrite_all($second, substr($pack, $second_pos));\n+\tclose($first) or die \"close first response: $!\";\n+\tclose($second) or die \"close second response: $!\";\n+\tEOF\n+\t{\n+\t\t(\n+\t\t\tif ! \"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n+\t\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n+\t\t\t\t\"$TRASH_DIRECTORY/first-ready\"\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/server-ready\" &&\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) >server.log 2>&1 &\n+\t\tserver_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $server_pid 2>/dev/null || :\n+\t\twait $server_pid 2>/dev/null || :\n+\t\texec 7>&-\n+\t\texec 8>&-\n+\t\trm -f server-ready first-ready slow-pack-server\n+\t\" &&\n+\tread port <&7 &&\n+\turl=\"http://127.0.0.1:$port/pack\" &&\n+\t{\n+\t\t(\n+\t\t\tif ! GIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-first.trace\" \\\n+\t\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$url\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\ttest_path_is_file \"$tmpfile\" &&\n+\ttest -s \"$tmpfile\" &&\n+\t{\n+\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n+\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\"$url\" >second.out &\n+\t\tsecond_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $second_pid 2>/dev/null || :\n+\t\twait $second_pid 2>/dev/null || :\n+\t\" &&\n+\twait \"$server_pid\" &&\n+\twait \"$first_pid\" &&\n+\twait \"$second_pid\" &&\n+\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n+\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n+\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n+\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n+\tsort first.out second.out >actual &&\n+\ttest_cmp expect actual &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-overlap cat-file -e \"$blob\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n-- \n2.55.0.125.g9b41d4ddb3\n\n"},{"id":"548702","messageId":"9b41d4ddb38a5d7d4cc84f5626d6a031155f8f05.1784582665.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784582665.git.tnyman@openai.com","subject":"[PATCH v2 2/2] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-20T22:34:00Z","receivedAt":"2026-07-20T22:34:07Z","isPatch":true,"body":"When index-pack finds an existing keep file it reports pack rather than\nkeep. Accept either result from http-fetch, and only register a keep\nlockfile when this fetch created it.\n\nRead the pack/keep prefix and hash without consuming any following fsck\noutput, validate the reported pack hash against the advertised hash, and\nexercise a packfile URI fetch with a pre-existing keep file.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n fetch-pack.c           | 33 ++++++++++++++++++---------------\n t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 15 deletions(-)\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 120e01f3cf..509b91527b 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t}\n \n \tfor (i = 0; i < packfile_uris.nr; i++) {\n+\t\tbool created_keep;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 1];\n+\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n-\t\t    memcmp(packname, \"keep\\t\", 5))\n-\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n+\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n+\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n+\t\t     memcmp(packhash, \"pack\\t\", 5)))\n+\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n+\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n \n-\t\tif (read_in_full(cmd.out, packname,\n-\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n-\t\t    packname[the_hash_algo->hexsz] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n-\n-\t\tpackname[the_hash_algo->hexsz] = '\\0';\n+\t\tif (read_in_full(cmd.out, packhash,\n+\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n+\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n+\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n+\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n \n \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n \n@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (finish_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to finish http-fetch\");\n \n-\t\tif (memcmp(packfile_uris.items[i].string, packname,\n+\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n \t\t\t   the_hash_algo->hexsz))\n \t\t\tdie(\"fetch-pack: pack downloaded from %s does not match expected hash %.*s\",\n \t\t\t    uri, (int) the_hash_algo->hexsz,\n \t\t\t    packfile_uris.items[i].string);\n \n-\t\tstring_list_append_nodup(pack_lockfiles,\n-\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n-\t\t\t\t\t\t repo_get_object_directory(the_repository),\n-\t\t\t\t\t\t packname));\n+\t\tif (created_keep)\n+\t\t\tstring_list_append_nodup(pack_lockfiles,\n+\t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n+\t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n+\t\t\t\t\t\t\t packhash));\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\ndiff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh\nindex 9f6cf4142d..1861eb7d7c 100755\n--- a/t/t5702-protocol-v2.sh\n+++ b/t/t5702-protocol-v2.sh\n@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '\n \t\tfetch \"$HTTPD_URL/smart/http_parent\"\n '\n \n+test_expect_success 'packfile URI preserves an existing keep file' '\n+\tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n+\trm -rf \"$P\" http_child keep.expect &&\n+\n+\tgit init \"$P\" &&\n+\tgit -C \"$P\" config uploadpack.allowsidebandall true &&\n+\n+\techo my-blob >\"$P/my-blob\" &&\n+\tgit -C \"$P\" add my-blob &&\n+\tgit -C \"$P\" commit -m x &&\n+\tconfigure_exclusion \"$P\" my-blob >h &&\n+\n+\tgit init http_child &&\n+\tpackhash=$(cat packh) &&\n+\tkeep=\"http_child/.git/objects/pack/pack-$packhash.keep\" &&\n+\techo pre-existing >\"$keep\" &&\n+\tcp \"$keep\" keep.expect &&\n+\n+\tGIT_TEST_SIDEBAND_ALL=1 \\\n+\tgit -C http_child -c protocol.version=2 \\\n+\t\t-c fetch.uriprotocols=http,https \\\n+\t\tfetch \"$HTTPD_URL/smart/http_parent\" &&\n+\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.pack\" &&\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.idx\" &&\n+\ttest_cmp keep.expect \"$keep\" &&\n+\tgit -C http_child cat-file -e \"$(cat h)\"\n+'\n+\n test_expect_success 'fetching with valid packfile URI but invalid hash fails' '\n \tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n \trm -rf \"$P\" http_child log &&\n-- \n2.55.0.125.g9b41d4ddb3\n\n"},{"id":"548736","messageId":"xmqqo6g0unfb.fsf@gitster.g","threadId":"65988","inReplyTo":"160a9b9fd0982dadfbf6f8fbb378d1a3e9173698.1784582665.git.tnyman@openai.com","subject":"Re: [PATCH v2 1/2] http: avoid concurrent appends to partial packs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-21T19:56:24Z","receivedAt":"2026-07-21T19:56:28Z","isPatch":true,"body":"Ted Nyman <tnyman@openai.com> writes:\n\n> Pack requests stage downloads in a predictable partial-pack file so an\n> interrupted transfer can be resumed. Both packfile URI and ordinary dumb\n> HTTP requests use this staging path. Opening it in append mode lets\n> concurrent fetches interleave their writes, corrupting the pack or\n> causing a later fetch to request a range at EOF.\n>\n> Open the partial pack read-write, seek to its current end, and retain a\n> per-descriptor offset for incoming data. Reopen newly created partial\n> packs without O_CREAT so Windows permits concurrent unlink, and keep the\n> descriptor for index-pack when another downloader removes the staging\n> path. Accept HTTP 416 when a partial pack is already complete.\n>\n> Exercise resumed transfers, EOF ranges, and overlapping 200 and 206\n> responses. Clarify the staging-key documentation and correct the stale\n> --index-pack-args spelling in the documentation and error messages; the\n> repeatable --index-pack-arg option is already accepted.\n\nHmph.  So the idea is to allow multiple processes to open the same\nfile and, because they all know where their respective chunks of\ndata fit in the final file, have them use pwrite(2) to deposit those\npieces at the exact target locations, and this prevents them from\nstepping on each other's toes?\n\nI cannot exactly explain why but it somehow makes me feel dirty.\n\nIt is also surprising that the workaround on MinGW works when\none of these multiple processes finishes writing and attempts to\nfinalize the temporary file while others still have open file\ndescriptors to the same file.\n\n> -\tThe hash is used to determine the name of the temporary file and is\n> -\tarbitrary. The output of index-pack is printed to stdout. Requires\n> -\t--index-pack-args.\n> +\tThe hash is used to determine the name of the temporary file. It need\n> +\tnot be the pack hash, but it must uniquely identify the pack contents\n> +\tfor resumption. The output of index-pack is printed to stdout. Requires\n> +\tone or more --index-pack-arg options.\n\nOK.\n\n> ---index-pack-args=<args>::\n> -\tFor internal use only. The command to run on the contents of the\n> -\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n> +--index-pack-arg=<arg>::\n> +\tFor internal use only. An argument to the command run on the contents\n> +\tof the downloaded pack. This option can be specified multiple times.\n\nWas the 'internal use only' thing renamed in order to prevent the\nnew code from accidentally working with an older caller?\n\n    ... goes and notices that the code uses singular form throughout ...\n\nAh, no, this is an unrelated typo fix that remains valid even if the\nrest of this patch is dropped.  Good catch.\n\nIt would be easier to review the actual changes if this cleanup were\nisolated in a preliminary patch.  Are there other cleanup changes in\nthis series that fall into the same category?\n\n> diff --git a/http-fetch.c b/http-fetch.c\n> index f9b6ecb061..05f68f306a 100644\n> --- a/http-fetch.c\n> +++ b/http-fetch.c\n> @@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,\n>  \n>  \tif (start_active_slot(preq->slot)) {\n>  \t\trun_active_slot(preq->slot);\n> -\t\tif (results.curl_result != CURLE_OK) {\n> +\t\tif (results.curl_result != CURLE_OK &&\n> +\t\t    results.http_code != 416) {\n\nWe do not seem to use symbolic constants for these '4xx' codes (or\n'2xx', for that matter), so I will let that pass.  Eventually, we\nmay want to give symbolic constants to them to improve readability,\nbut doing so is certainly outside the scope of this topic.\n\n> @@ -155,7 +156,7 @@ int cmd_main(int argc, const char **argv)\n>  \n>  \tif (packfile) {\n>  \t\tif (!index_pack_args.nr)\n> -\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n> +\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n\nThis and ...\n\n> @@ -164,7 +165,7 @@ int cmd_main(int argc, const char **argv)\n>  \t}\n>  \n>  \tif (index_pack_args.nr)\n> -\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n> +\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n>  \n>  \tif (commits_on_stdin) {\n>  \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n\n... this is the same \"index-pack-arg\" fix and can be moved to a\nseparate preliminary clean-up patch.\n\nThanks.\n\n"},{"id":"548749","messageId":"cover.1784676106.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1783982021.git.tnyman@openai.com","subject":"[PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-21T23:29:39Z","receivedAt":"2026-07-21T23:30:12Z","isPatch":true,"body":"Packfile URI and dumb HTTP downloads stage packs at\nobjects/pack/pack-<hash>.pack.temp so an interrupted transfer can\nresume. Opening that file in append mode forces every write to its\ncurrent end. Two Git processes fetching the same pack into one object\ndatabase can therefore append duplicate data and corrupt the pack.\n\nThe first patch separates the unrelated --index-pack-arg documentation\nand error-message correction requested during review.\n\nThe second patch keeps the predictable staging name but removes append\nmode. Each downloader seeks once to the current end, requests the\ncorresponding Range, and writes using its own descriptor offset. Since\nthe staging key must identify immutable pack contents, overlapping\nresponses write identical bytes at identical offsets. There is no need\nfor pwrite(2) or cross-process coordination, and resumption continues to\nwork for both packfile URI and ordinary dumb HTTP downloads.\n\nA downloader can also find that the partial pack has completed and\nrequest a range starting at EOF. Servers may respond with HTTP 416 in\nthat case. Treat the response as a completed download and let\nindex-pack validate the pack.\n\nOn MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an\nexisting file. Create a missing staging file exclusively, close it, and\nreopen it without O_CREAT so every retained descriptor permits another\ndownloader to unlink the path. Keep the open descriptor for index-pack;\nit installs its own pack, so the shared staging file is only unlinked,\nnever renamed.\n\nThe third patch handles the related .keep race. When another process has\nalready created the keep file, index-pack reports \"pack<TAB><hash>\"\ninstead of \"keep<TAB><hash>\". Accept both successful forms and remove\nonly keep files created by the current process. Read only the prefix and\nhash so any following fsck output remains available to fetch-pack.\n\nThe tests cover resumption, a completed partial returning 416,\noverlapping 200 and 206 responses, unlinking the staging path while\nindex-pack holds its descriptor, and a pre-existing .keep file. The\nunlink test does not require FIFOs, so it can exercise MinGW's sharing\nbehavior even though the concurrent-download tests are skipped there.\n\nChanges since v2:\n\n  * Split the --index-pack-arg documentation and error-message cleanup\n    into a preliminary patch, as requested by Junio.\n  * Clarify why per-descriptor offsets keep overlapping writes safe and\n    why MinGW permits the shared staging path to be unlinked.\n  * Add a non-FIFO unlink-while-indexing regression test that can run on\n    MinGW.\n  * Rebase onto the current master.\n\nThe v2 discussion is at:\n\n  https://lore.kernel.org/git/cover.1784582665.git.tnyman@openai.com/\n\nTed Nyman (3):\n  http-fetch: correct --index-pack-arg documentation\n  http: avoid concurrent appends to partial packs\n  fetch-pack: accept \"pack\" output for packfile URIs\n\n Documentation/git-http-fetch.adoc |  13 +-\n fetch-pack.c                      |  33 ++--\n http-fetch.c                      |   7 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 ++++---\n t/t5550-http-fetch-dumb.sh        | 244 ++++++++++++++++++++++++++++++\n t/t5702-protocol-v2.sh            |  31 ++++\n 8 files changed, 344 insertions(+), 46 deletions(-)\n\nRange-diff against v2:\n-:  ---------- > 1:  a6a40b8046 http-fetch: correct --index-pack-arg documentation\n1:  160a9b9fd0 ! 2:  6c91054afc http: avoid concurrent appends to partial packs\n    @@ Commit message\n     \n         Pack requests stage downloads in a predictable partial-pack file so an\n         interrupted transfer can be resumed. Both packfile URI and ordinary dumb\n    -    HTTP requests use this staging path. Opening it in append mode lets\n    -    concurrent fetches interleave their writes, corrupting the pack or\n    -    causing a later fetch to request a range at EOF.\n    +    HTTP requests use this staging path. Opening it in append mode forces\n    +    each write to the current end of the file, so concurrent responses can\n    +    append duplicate data and corrupt the pack.\n     \n    -    Open the partial pack read-write, seek to its current end, and retain a\n    -    per-descriptor offset for incoming data. Reopen newly created partial\n    -    packs without O_CREAT so Windows permits concurrent unlink, and keep the\n    -    descriptor for index-pack when another downloader removes the staging\n    -    path. Accept HTTP 416 when a partial pack is already complete.\n    +    Open the partial pack read-write without O_APPEND and seek once to its\n    +    current end. Each downloader then retains the offset matching the Range\n    +    it requested. Because the staging key must uniquely identify immutable\n    +    pack contents, overlapping responses write the same bytes at the same\n    +    offsets instead of extending the file with duplicate data.\n     \n    -    Exercise resumed transfers, EOF ranges, and overlapping 200 and 206\n    -    responses. Clarify the staging-key documentation and correct the stale\n    -    --index-pack-args spelling in the documentation and error messages; the\n    -    repeatable --index-pack-arg option is already accepted.\n    +    MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n    +    existing file. Create a missing partial pack exclusively, close it, and\n    +    reopen it without O_CREAT so every retained descriptor permits another\n    +    downloader to unlink the staging path. Duplicate that descriptor for\n    +    index-pack instead of reopening the path after closing the stream;\n    +    index-pack installs its own pack and the shared staging file is only\n    +    unlinked, never renamed. Accept HTTP 416 when a partial pack is already\n    +    complete and let index-pack validate its contents.\n    +\n    +    Exercise resumed transfers, EOF ranges, overlapping 200 and 206\n    +    responses, and unlinking the staging path while index-pack still holds\n    +    its descriptor. Clarify the staging-key documentation.\n     \n         Signed-off-by: Ted Nyman <tnyman@openai.com>\n     \n    @@ Documentation/git-http-fetch.adoc: commit-id::\n      \tURL and uses index-pack to generate corresponding .idx and .keep files.\n     -\tThe hash is used to determine the name of the temporary file and is\n     -\tarbitrary. The output of index-pack is printed to stdout. Requires\n    --\t--index-pack-args.\n     +\tThe hash is used to determine the name of the temporary file. It need\n     +\tnot be the pack hash, but it must uniquely identify the pack contents\n     +\tfor resumption. The output of index-pack is printed to stdout. Requires\n    -+\tone or more --index-pack-arg options.\n    - \n    ----index-pack-args=<args>::\n    --\tFor internal use only. The command to run on the contents of the\n    --\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n    -+--index-pack-arg=<arg>::\n    -+\tFor internal use only. An argument to the command run on the contents\n    -+\tof the downloaded pack. This option can be specified multiple times.\n    + \tone or more --index-pack-arg options.\n      \n    - --recover::\n    - \tVerify that everything reachable from target is fetched.  Used after\n    + --index-pack-arg=<arg>::\n     \n      ## http-fetch.c ##\n     @@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,\n    @@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,\n      \t\t\tstruct url_info url;\n      \t\t\tchar *nurl = url_normalize(preq->url, &url);\n      \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\n    -@@ http-fetch.c: int cmd_main(int argc, const char **argv)\n    - \n    - \tif (packfile) {\n    - \t\tif (!index_pack_args.nr)\n    --\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n    -+\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n    - \n    - \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n    - \t\t\t\t      index_pack_args.v);\n    -@@ http-fetch.c: int cmd_main(int argc, const char **argv)\n    - \t}\n    - \n    - \tif (index_pack_args.nr)\n    --\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n    -+\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n    - \n    - \tif (commits_on_stdin) {\n    - \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n     \n      ## http-push.c ##\n     @@ http-push.c: static void finish_request(struct transfer_request *request)\n    @@ http.c: struct http_pack_request *new_http_pack_request(\n     -\tif (!preq->packfile) {\n     -\t\terror(\"Unable to open local file %s for pack\",\n     -\t\t      preq->tmpfile.buf);\n    -+\t/* Reopen without O_CREAT so MinGW permits another writer to unlink it. */\n    ++\t/*\n    ++\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n    ++\t * existing file; reopen a newly created file so others may unlink it.\n    ++\t */\n     +\tfor (;;) {\n     +\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n     +\t\tif (fd >= 0 || errno != ENOENT)\n    @@ http.c: struct http_pack_request *new_http_pack_request(\n     +\tif (fd < 0) {\n     +\t\terror_errno(\"unable to open local file %s for pack\",\n     +\t\t\t    preq->tmpfile.buf);\n    -+\t\tgoto abort;\n    -+\t}\n    + \t\tgoto abort;\n    + \t}\n     +\tprev_posn = lseek(fd, 0, SEEK_END);\n     +\tif (prev_posn < 0) {\n     +\t\terror_errno(\"unable to seek local file %s for pack\",\n     +\t\t\t    preq->tmpfile.buf);\n     +\t\tclose(fd);\n    - \t\tgoto abort;\n    - \t}\n    ++\t\tgoto abort;\n    ++\t}\n     +\tpreq->packfile = xfdopen(fd, \"w\");\n      \n      \tpreq->slot = get_active_slot();\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n     +'\n     +\n    ++test_expect_success 'http-fetch --packfile permits unlink while indexing' '\n    ++\tgit init packfileclient-unlink &&\n    ++\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n    ++\t\tls objects/pack/pack-*.pack) &&\n    ++\ttmpfile=\"packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n    ++\twrite_script git-unlink-index-pack <<-\\EOF &&\n    ++\ttest -f \"$GIT_TEST_PACK_TEMP\" || exit 1\n    ++\trm \"$GIT_TEST_PACK_TEMP\" || exit 1\n    ++\texec git index-pack \"$@\"\n    ++\tEOF\n    ++\ttest_when_finished \"rm -f git-unlink-index-pack\" &&\n    ++\tPATH=\"$TRASH_DIRECTORY:$PATH\" \\\n    ++\tGIT_TEST_PACK_TEMP=\"$TRASH_DIRECTORY/$tmpfile\" \\\n    ++\tgit -C packfileclient-unlink http-fetch --packfile=\"$ARBITRARY\" \\\n    ++\t\t--index-pack-arg=unlink-index-pack \\\n    ++\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n    ++\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n    ++\ttest_path_is_missing \"$tmpfile\" &&\n    ++\tgit -C packfileclient-unlink cat-file -e \"$HASH\"\n    ++'\n    ++\n     +test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n     +\tgit init packfileclient-concurrent &&\n     +\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n2:  9b41d4ddb3 = 3:  1ee5d7e027 fetch-pack: accept \"pack\" output for packfile URIs\n\nbase-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200\n-- \n2.55.0.openai.131.g83a728de1eb6\n"},{"id":"548751","messageId":"a6a40b80461377452a0b2c9204c3a659ab60a7d5.1784676106.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784676106.git.tnyman@openai.com","subject":"[PATCH v3 1/3] http-fetch: correct --index-pack-arg documentation","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-21T23:29:40Z","receivedAt":"2026-07-21T23:30:14Z","isPatch":true,"body":"The --packfile mode accepts one --index-pack-arg=<arg> option per\nargument passed to index-pack, but its documentation and option\ndependency errors still refer to the plural --index-pack-args form.\n\nCorrect the spelling and describe the repeatable per-argument form.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc | 8 ++++----\n http-fetch.c                      | 4 ++--\n 2 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 2200f073c4..09b5d675ee 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -50,11 +50,11 @@ commit-id::\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n \tThe hash is used to determine the name of the temporary file and is\n \tarbitrary. The output of index-pack is printed to stdout. Requires\n-\t--index-pack-args.\n+\tone or more --index-pack-arg options.\n \n---index-pack-args=<args>::\n-\tFor internal use only. The command to run on the contents of the\n-\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n+--index-pack-arg=<arg>::\n+\tFor internal use only. An argument to the command run on the contents\n+\tof the downloaded pack. This option can be specified multiple times.\n \n --recover::\n \tVerify that everything reachable from target is fetched.  Used after\ndiff --git a/http-fetch.c b/http-fetch.c\nindex f9b6ecb061..601a77c3c1 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)\n \n \tif (packfile) {\n \t\tif (!index_pack_args.nr)\n-\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n+\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n \n \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n \t\t\t\t      index_pack_args.v);\n@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)\n \t}\n \n \tif (index_pack_args.nr)\n-\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n+\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n \n \tif (commits_on_stdin) {\n \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548750","messageId":"6c91054afcf911f10450df036526d7d374e1e56f.1784676106.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784676106.git.tnyman@openai.com","subject":"[PATCH v3 2/3] http: avoid concurrent appends to partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-21T23:29:41Z","receivedAt":"2026-07-21T23:30:16Z","isPatch":true,"body":"Pack requests stage downloads in a predictable partial-pack file so an\ninterrupted transfer can be resumed. Both packfile URI and ordinary dumb\nHTTP requests use this staging path. Opening it in append mode forces\neach write to the current end of the file, so concurrent responses can\nappend duplicate data and corrupt the pack.\n\nOpen the partial pack read-write without O_APPEND and seek once to its\ncurrent end. Each downloader then retains the offset matching the Range\nit requested. Because the staging key must uniquely identify immutable\npack contents, overlapping responses write the same bytes at the same\noffsets instead of extending the file with duplicate data.\n\nMinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\nexisting file. Create a missing partial pack exclusively, close it, and\nreopen it without O_CREAT so every retained descriptor permits another\ndownloader to unlink the staging path. Duplicate that descriptor for\nindex-pack instead of reopening the path after closing the stream;\nindex-pack installs its own pack and the shared staging file is only\nunlinked, never renamed. Accept HTTP 416 when a partial pack is already\ncomplete and let index-pack validate its contents.\n\nExercise resumed transfers, EOF ranges, overlapping 200 and 206\nresponses, and unlinking the staging path while index-pack still holds\nits descriptor. Clarify the staging-key documentation.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc |   5 +-\n http-fetch.c                      |   3 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 ++++---\n t/t5550-http-fetch-dumb.sh        | 244 ++++++++++++++++++++++++++++++\n 6 files changed, 289 insertions(+), 25 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 09b5d675ee..60ca91cf3a 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -48,8 +48,9 @@ commit-id::\n \tline (which is not expected in\n \tthis case), 'git http-fetch' fetches the packfile directly at the given\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n-\tThe hash is used to determine the name of the temporary file and is\n-\tarbitrary. The output of index-pack is printed to stdout. Requires\n+\tThe hash is used to determine the name of the temporary file. It need\n+\tnot be the pack hash, but it must uniquely identify the pack contents\n+\tfor resumption. The output of index-pack is printed to stdout. Requires\n \tone or more --index-pack-arg options.\n \n --index-pack-arg=<arg>::\ndiff --git a/http-fetch.c b/http-fetch.c\nindex 601a77c3c1..05f68f306a 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\tstruct url_info url;\n \t\t\tchar *nurl = url_normalize(preq->url, &url);\n \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\ndiff --git a/http-push.c b/http-push.c\nindex 60f6f8f054..ef8abe3908 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)\n \n \t} else if (request->state == RUN_FETCH_PACKED) {\n \t\tint fail = 1;\n-\t\tif (request->curl_result != CURLE_OK) {\n+\t\tif (request->curl_result != CURLE_OK &&\n+\t\t    request->http_code != 416) {\n \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n \t\t\t\trequest->url, curl_errorstr);\n \t\t} else {\ndiff --git a/http-walker.c b/http-walker.c\nindex b58a3b2a92..abafca84d6 100644\n--- a/http-walker.c\n+++ b/http-walker.c\n@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n \t\t\t      curl_errorstr);\n \t\t\tgoto abort;\ndiff --git a/http.c b/http.c\nindex caccf2108e..a0d399b274 100644\n--- a/http.c\n+++ b/http.c\n@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \tint tmpfile_fd;\n \tint ret = 0;\n \n+\t/* Another downloader may unlink the staging path while we index it. */\n+\ttmpfile_fd = xdup(fileno(preq->packfile));\n \tfclose(preq->packfile);\n \tpreq->packfile = NULL;\n-\n-\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n+\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n+\t\tdie_errno(\"unable to seek local file %s for pack\",\n+\t\t\t  preq->tmpfile.buf);\n \n \tip.git_cmd = 1;\n \tip.in = tmpfile_fd;\n@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \telse\n \t\tip.no_stdout = 1;\n \n-\tif (run_command(&ip)) {\n+\tif (run_command(&ip))\n \t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\n-cleanup:\n-\tclose(tmpfile_fd);\n \tunlink(preq->tmpfile.buf);\n \treturn ret;\n }\n@@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(\n struct http_pack_request *new_direct_http_pack_request(\n \tconst unsigned char *packed_git_hash, char *url)\n {\n-\toff_t prev_posn = 0;\n+\toff_t prev_posn;\n \tstruct http_pack_request *preq;\n+\tint fd;\n \n \tCALLOC_ARRAY(preq, 1);\n \tstrbuf_init(&preq->tmpfile, 0);\n-\n \tpreq->url = url;\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n-\tif (!preq->packfile) {\n-\t\terror(\"Unable to open local file %s for pack\",\n-\t\t      preq->tmpfile.buf);\n+\t/*\n+\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n+\t * existing file; reopen a newly created file so others may unlink it.\n+\t */\n+\tfor (;;) {\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n+\t\tif (fd >= 0 || errno != ENOENT)\n+\t\t\tbreak;\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n+\t\tif (fd >= 0) {\n+\t\t\tclose(fd);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (errno != EEXIST)\n+\t\t\tbreak;\n+\t}\n+\tif (fd < 0) {\n+\t\terror_errno(\"unable to open local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n \t\tgoto abort;\n \t}\n+\tprev_posn = lseek(fd, 0, SEEK_END);\n+\tif (prev_posn < 0) {\n+\t\terror_errno(\"unable to seek local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tclose(fd);\n+\t\tgoto abort;\n+\t}\n+\tpreq->packfile = xfdopen(fd, \"w\");\n \n \tpreq->slot = get_active_slot();\n \tpreq->headers = object_request_headers();\n@@ -2762,12 +2783,7 @@ struct http_pack_request *new_direct_http_pack_request(\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);\n \n-\t/*\n-\t * If there is data present from a previous transfer attempt,\n-\t * resume where it left off\n-\t */\n-\tprev_posn = ftello(preq->packfile);\n-\tif (prev_posn>0) {\n+\tif (prev_posn > 0) {\n \t\tif (http_is_verbose)\n \t\t\tfprintf(stderr,\n \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex f00eeae48f..65b42c4719 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -293,6 +293,250 @@ test_expect_success 'http-fetch --packfile' '\n \tgit -C packfileclient cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile resumes a partial download' '\n+\tgit init packfileclient-resume &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/resume.trace\" \\\n+\tgit -C packfileclient-resume http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=index-pack --index-pack-arg=--stdin \\\n+\t\t--index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_grep \"Range: bytes=64-\" resume.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success 'http-fetch --packfile permits unlink while indexing' '\n+\tgit init packfileclient-unlink &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\twrite_script git-unlink-index-pack <<-\\EOF &&\n+\ttest -f \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\trm \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\texec git index-pack \"$@\"\n+\tEOF\n+\ttest_when_finished \"rm -f git-unlink-index-pack\" &&\n+\tPATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\tGIT_TEST_PACK_TEMP=\"$TRASH_DIRECTORY/$tmpfile\" \\\n+\tgit -C packfileclient-unlink http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=unlink-index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-unlink cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n+\tgit init packfileclient-concurrent &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\tpackhash=$(basename \"$p\" .pack) &&\n+\tpackhash=${packhash#pack-} &&\n+\ttmpfile=\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tmkfifo first-ready first-continue &&\n+\texec 8<>first-ready &&\n+\texec 9<>first-continue &&\n+\twrite_script git-wait-index-pack <<-\\EOF &&\n+\techo ready >\"$GIT_TEST_WAIT_READY\" &&\n+\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n+\texec git index-pack \"$@\"\n+\tEOF\n+\t{\n+\t\t(\n+\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n+\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n+\t\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/first.trace\" \\\n+\t\t\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=wait-index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\techo continue >&9\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\texec 8>&-\n+\t\texec 9>&-\n+\t\trm -f first-ready first-continue git-wait-index-pack\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/second.trace\" \\\n+\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t--index-pack-arg=index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n+\techo continue >&9 &&\n+\twait \"$first_pid\" &&\n+\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect first.out &&\n+\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect second.out &&\n+\ttest_grep \"Range: bytes=64-\" first.trace &&\n+\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n+\ttest_grep \"HTTP/[0-9.]* 416\" second.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n+\tgit init packfileclient-overlap &&\n+\tblob=$(test-tool genrandom pack-overlap 2m |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\thash-object -w --stdin) &&\n+\tpackhash=$(printf \"%s\\n\" \"$blob\" |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\tpack-objects \"$TRASH_DIRECTORY/overlap-pack\") &&\n+\tpack=\"$TRASH_DIRECTORY/overlap-pack-$packhash.pack\" &&\n+\ttmpfile=\"packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tmkfifo server-ready first-ready &&\n+\texec 7<>server-ready &&\n+\texec 8<>first-ready &&\n+\twrite_script slow-pack-server \"$PERL_PATH\" <<-\\EOF &&\n+\tuse strict;\n+\tuse warnings;\n+\tuse IO::Socket::INET;\n+\n+\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n+\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n+\tmy $pack = do { local $/; <$in> };\n+\tclose($in) or die \"close $packfile: $!\";\n+\tmy $server = IO::Socket::INET->new(LocalAddr => \"127.0.0.1\",\n+\t\tLocalPort => 0, Proto => \"tcp\", Listen => 2, ReuseAddr => 1)\n+\t\tor die \"listen: $!\";\n+\n+\tsub signal_ready {\n+\t\tmy ($file, $value) = @_;\n+\t\topen(my $out, \">\", $file) or die \"open $file: $!\";\n+\t\tprint $out \"$value\\n\" or die \"write $file: $!\";\n+\t\tclose($out) or die \"close $file: $!\";\n+\t}\n+\n+\tsub write_all {\n+\t\tmy ($out, $data) = @_;\n+\t\tmy $offset = 0;\n+\t\twhile ($offset < length($data)) {\n+\t\t\tmy $written = syswrite($out, $data,\n+\t\t\t\tlength($data) - $offset, $offset);\n+\t\t\tdefined($written) && $written or die \"write response: $!\";\n+\t\t\t$offset += $written;\n+\t\t}\n+\t}\n+\n+\tsub start_response {\n+\t\tmy $out = $server->accept() or die \"accept: $!\";\n+\t\t<$out> or die \"read request: $!\";\n+\t\tmy $start = 0;\n+\t\twhile (<$out>) {\n+\t\t\tlast if /^\\r?\\n$/;\n+\t\t\t$start = $1 if /^Range: bytes=(\\d+)-/i;\n+\t\t}\n+\t\t$start < length($pack) or die \"invalid range $start\";\n+\t\tmy $length = length($pack) - $start;\n+\t\tmy $middle = int($length / 2);\n+\t\tmy $status = $start ? \"206 Partial Content\" : \"200 OK\";\n+\t\tmy $headers = \"HTTP/1.1 $status\\r\\n\" .\n+\t\t\t\"Content-Length: $length\\r\\n\" .\n+\t\t\t($start ? \"Content-Range: bytes $start-\" .\n+\t\t\t\t(length($pack) - 1) . \"/\" . length($pack) . \"\\r\\n\" : \"\") .\n+\t\t\t\"Connection: close\\r\\n\\r\\n\";\n+\t\twrite_all($out, $headers);\n+\t\twrite_all($out, substr($pack, $start, $middle));\n+\t\treturn ($out, $start + $middle);\n+\t}\n+\n+\tsignal_ready($server_ready, $server->sockport());\n+\tmy ($first, $first_pos) = start_response();\n+\tsignal_ready($first_ready, \"ready\");\n+\tmy ($second, $second_pos) = start_response();\n+\twrite_all($first, substr($pack, $first_pos));\n+\twrite_all($second, substr($pack, $second_pos));\n+\tclose($first) or die \"close first response: $!\";\n+\tclose($second) or die \"close second response: $!\";\n+\tEOF\n+\t{\n+\t\t(\n+\t\t\tif ! \"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n+\t\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n+\t\t\t\t\"$TRASH_DIRECTORY/first-ready\"\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/server-ready\" &&\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) >server.log 2>&1 &\n+\t\tserver_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $server_pid 2>/dev/null || :\n+\t\twait $server_pid 2>/dev/null || :\n+\t\texec 7>&-\n+\t\texec 8>&-\n+\t\trm -f server-ready first-ready slow-pack-server\n+\t\" &&\n+\tread port <&7 &&\n+\turl=\"http://127.0.0.1:$port/pack\" &&\n+\t{\n+\t\t(\n+\t\t\tif ! GIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-first.trace\" \\\n+\t\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$url\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\ttest_path_is_file \"$tmpfile\" &&\n+\ttest -s \"$tmpfile\" &&\n+\t{\n+\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n+\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\"$url\" >second.out &\n+\t\tsecond_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $second_pid 2>/dev/null || :\n+\t\twait $second_pid 2>/dev/null || :\n+\t\" &&\n+\twait \"$server_pid\" &&\n+\twait \"$first_pid\" &&\n+\twait \"$second_pid\" &&\n+\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n+\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n+\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n+\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n+\tsort first.out second.out >actual &&\n+\ttest_cmp expect actual &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-overlap cat-file -e \"$blob\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548752","messageId":"1ee5d7e02747e76dd044571379ce07bdc9c96c94.1784676106.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784676106.git.tnyman@openai.com","subject":"[PATCH v3 3/3] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-21T23:29:42Z","receivedAt":"2026-07-21T23:30:18Z","isPatch":true,"body":"When index-pack finds an existing keep file it reports pack rather than\nkeep. Accept either result from http-fetch, and only register a keep\nlockfile when this fetch created it.\n\nRead the pack/keep prefix and hash without consuming any following fsck\noutput, validate the reported pack hash against the advertised hash, and\nexercise a packfile URI fetch with a pre-existing keep file.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n fetch-pack.c           | 33 ++++++++++++++++++---------------\n t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 15 deletions(-)\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 29c41132ee..e9f24fbd63 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t}\n \n \tfor (i = 0; i < packfile_uris.nr; i++) {\n+\t\tbool created_keep;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 1];\n+\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n-\t\t    memcmp(packname, \"keep\\t\", 5))\n-\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n+\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n+\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n+\t\t     memcmp(packhash, \"pack\\t\", 5)))\n+\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n+\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n \n-\t\tif (read_in_full(cmd.out, packname,\n-\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n-\t\t    packname[the_hash_algo->hexsz] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n-\n-\t\tpackname[the_hash_algo->hexsz] = '\\0';\n+\t\tif (read_in_full(cmd.out, packhash,\n+\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n+\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n+\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n+\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n \n \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n \n@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (finish_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to finish http-fetch\");\n \n-\t\tif (memcmp(packfile_uris.items[i].string, packname,\n+\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n \t\t\t   the_hash_algo->hexsz))\n \t\t\tdie(\"fetch-pack: pack downloaded from %s does not match expected hash %.*s\",\n \t\t\t    uri, (int) the_hash_algo->hexsz,\n \t\t\t    packfile_uris.items[i].string);\n \n-\t\tstring_list_append_nodup(pack_lockfiles,\n-\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n-\t\t\t\t\t\t repo_get_object_directory(the_repository),\n-\t\t\t\t\t\t packname));\n+\t\tif (created_keep)\n+\t\t\tstring_list_append_nodup(pack_lockfiles,\n+\t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n+\t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n+\t\t\t\t\t\t\t packhash));\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\ndiff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh\nindex 74a2b7730b..0f05286de8 100755\n--- a/t/t5702-protocol-v2.sh\n+++ b/t/t5702-protocol-v2.sh\n@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '\n \t\tfetch \"$HTTPD_URL/smart/http_parent\"\n '\n \n+test_expect_success 'packfile URI preserves an existing keep file' '\n+\tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n+\trm -rf \"$P\" http_child keep.expect &&\n+\n+\tgit init \"$P\" &&\n+\tgit -C \"$P\" config uploadpack.allowsidebandall true &&\n+\n+\techo my-blob >\"$P/my-blob\" &&\n+\tgit -C \"$P\" add my-blob &&\n+\tgit -C \"$P\" commit -m x &&\n+\tconfigure_exclusion \"$P\" my-blob >h &&\n+\n+\tgit init http_child &&\n+\tpackhash=$(cat packh) &&\n+\tkeep=\"http_child/.git/objects/pack/pack-$packhash.keep\" &&\n+\techo pre-existing >\"$keep\" &&\n+\tcp \"$keep\" keep.expect &&\n+\n+\tGIT_TEST_SIDEBAND_ALL=1 \\\n+\tgit -C http_child -c protocol.version=2 \\\n+\t\t-c fetch.uriprotocols=http,https \\\n+\t\tfetch \"$HTTPD_URL/smart/http_parent\" &&\n+\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.pack\" &&\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.idx\" &&\n+\ttest_cmp keep.expect \"$keep\" &&\n+\tgit -C http_child cat-file -e \"$(cat h)\"\n+'\n+\n test_expect_success 'fetching with valid packfile URI but invalid hash fails' '\n \tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n \trm -rf \"$P\" http_child log &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548860","messageId":"xmqqldb19evx.fsf@gitster.g","threadId":"65988","inReplyTo":"cover.1784676106.git.tnyman@openai.com","subject":"Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-24T04:43:14Z","receivedAt":"2026-07-24T04:43:17Z","isPatch":true,"body":"Ted Nyman <tnyman@openai.com> writes:\n\n> Packfile URI and dumb HTTP downloads stage packs at\n> objects/pack/pack-<hash>.pack.temp so an interrupted transfer can\n> resume. Opening that file in append mode forces every write to its\n> current end. Two Git processes fetching the same pack into one object\n> database can therefore append duplicate data and corrupt the pack.\n> ...\n> The tests cover resumption, a completed partial returning 416,\n> overlapping 200 and 206 responses, unlinking the staging path while\n> index-pack holds its descriptor, and a pre-existing .keep file. The\n> unlink test does not require FIFOs, so it can exercise MinGW's sharing\n> behavior even though the concurrent-download tests are skipped there.\n>\n> Changes since v2:\n>\n>   * Split the --index-pack-arg documentation and error-message cleanup\n>     into a preliminary patch, as requested by Junio.\n>   * Clarify why per-descriptor offsets keep overlapping writes safe and\n>     why MinGW permits the shared staging path to be unlinked.\n>   * Add a non-FIFO unlink-while-indexing regression test that can run on\n>     MinGW.\n>   * Rebase onto the current master.\n\nWhen merged into 'seen', this topic seems to cause t5550 to hang\nfairly consistently.  It is not surprising, considering that the\ntopic adds roughly 240 lines to the test script in question.  It is\nentirely possible that we are seeing an existing breakage from\nanother topic in 'seen' that is exposed by the additional tests.\n\nThe CI run\n\n  https://github.com/git/git/actions/runs/30045343889\n\nis today's seen (excluding this topic) at 728e180b7b; it has\nbreakages in leak checking jobs from other topics, but does not see\nt5550 hanging.\n\nThe CI run\n\n  https://github.com/git/git/actions/runs/30048327878\n\nis seen at 05d0dd408c that merges this topic on top of 728e180b7b\nabove.  It breaks the same leak checks, but in addition makes t5550\nhang.\n\nCan you help figure out what is going on?\n\nThanks.\n\n\nPS. Recent CI runs on 'seen' started to spend so much time on static\n    analysis (aka coccinelle) jobs, even though I do not think we\n    acquired any new rules recently.  We probably need to figure out\n    what is going on there, too.  There is something wrong for these\n    CI runs that usually take ~40 minutes to spin for more than 4\n    hours.\n\n"},{"id":"548861","messageId":"cover.1784874850.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784676106.git.tnyman@openai.com","subject":"[PATCH v4 0/3] packfile URIs: support concurrent downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-24T08:14:22Z","receivedAt":"2026-07-24T08:14:24Z","isPatch":true,"body":"Packfile URI and dumb HTTP downloads stage packs at\nobjects/pack/pack-<hash>.pack.temp so an interrupted transfer can\nresume. Opening that file in append mode forces every write to its\ncurrent end. Two Git processes fetching the same pack into one object\ndatabase can therefore append duplicate data and corrupt the pack.\n\nThe first patch separates the unrelated --index-pack-arg documentation\nand error-message correction requested during review.\n\nThe second patch keeps the predictable staging name but removes append\nmode. Each downloader seeks once to the current end, requests the\ncorresponding Range, and writes using its own descriptor offset. Since\nthe staging key must identify immutable pack contents, overlapping\nresponses write identical bytes at identical offsets. There is no need\nfor pwrite(2) or cross-process coordination, and resumption continues to\nwork for both packfile URI and ordinary dumb HTTP downloads.\n\nA downloader can also find that the partial pack has completed and\nrequest a range starting at EOF. Servers may respond with HTTP 416 in\nthat case. Treat the response as a completed download and let\nindex-pack validate the pack.\n\nOn MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an\nexisting file. Create a missing staging file exclusively, close it, and\nreopen it without O_CREAT so every retained descriptor permits another\ndownloader to unlink the path. Keep the open descriptor for index-pack;\nit installs its own pack, so the shared staging file is only unlinked,\nnever renamed.\n\nThe third patch handles the related .keep race. When another process has\nalready created the keep file, index-pack reports \"pack<TAB><hash>\"\ninstead of \"keep<TAB><hash>\". Accept both successful forms and remove\nonly keep files created by the current process. Read only the prefix and\nhash so any following fsck output remains available to fetch-pack.\n\nThe tests cover resumption, a completed partial returning 416,\noverlapping 200 and 206 responses, unlinking the staging path while\nindex-pack holds its descriptor, and a pre-existing .keep file. The\nunlink test does not require FIFOs, so it can exercise MinGW's sharing\nbehavior even though the concurrent-download tests are skipped there.\n\nChanges since v3:\n\n  * Match HTTP 416 in trace output from both older and current libcurl.\n  * Add a timeout to the overlapping-download test server, notify FIFO\n    waiters on server failures, and track the actual server process for\n    cleanup.\n  * Wait for the second downloader first so an early failure cannot\n    leave the test server waiting for a request that will never arrive.\n  * No production code changes.\n\nThese changes avoid false failures with older libcurl and prevent a\nfailed downloader from leaving the test server running indefinitely.\n\nThe v3 discussion is at:\n\n  https://lore.kernel.org/git/cover.1784676106.git.tnyman@openai.com/\n\nTed Nyman (3):\n  http-fetch: correct --index-pack-arg documentation\n  http: avoid concurrent appends to partial packs\n  fetch-pack: accept \"pack\" output for packfile URIs\n\n Documentation/git-http-fetch.adoc |  13 +-\n fetch-pack.c                      |  33 ++--\n http-fetch.c                      |   7 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 ++++---\n t/t5550-http-fetch-dumb.sh        | 250 ++++++++++++++++++++++++++++++\n t/t5702-protocol-v2.sh            |  31 ++++\n 8 files changed, 350 insertions(+), 46 deletions(-)\n\nRange-diff against v3:\n1:  a6a40b8046 = 1:  a6a40b8046 http-fetch: correct --index-pack-arg documentation\n2:  6c91054afc ! 2:  144c98cdfa http: avoid concurrent appends to partial packs\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\ttest_cmp expect second.out &&\n     +\ttest_grep \"Range: bytes=64-\" first.trace &&\n     +\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n    -+\ttest_grep \"HTTP/[0-9.]* 416\" second.trace &&\n    ++\ttest_grep \"416 Requested Range Not Satisfiable\" second.trace &&\n     +\ttest_path_is_missing \"$tmpfile\" &&\n     +\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n     +'\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\tuse IO::Socket::INET;\n     +\n     +\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n    ++\tmy $completed = 0;\n    ++\tEND {\n    ++\t\tif (!$completed) {\n    ++\t\t\tsignal_ready($server_ready, \"failed\");\n    ++\t\t\tsignal_ready($first_ready, \"failed\");\n    ++\t\t}\n    ++\t}\n    ++\n    ++\t$SIG{ALRM} = sub { die \"timed out serving concurrent pack requests\\n\" };\n    ++\talarm 60;\n    ++\n     +\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n     +\tmy $pack = do { local $/; <$in> };\n     +\tclose($in) or die \"close $packfile: $!\";\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\twrite_all($second, substr($pack, $second_pos));\n     +\tclose($first) or die \"close first response: $!\";\n     +\tclose($second) or die \"close second response: $!\";\n    ++\t$completed = 1;\n    ++\talarm 0;\n     +\tEOF\n     +\t{\n    -+\t\t(\n    -+\t\t\tif ! \"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n    -+\t\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n    -+\t\t\t\t\"$TRASH_DIRECTORY/first-ready\"\n    -+\t\t\tthen\n    -+\t\t\t\techo failed >\"$TRASH_DIRECTORY/server-ready\" &&\n    -+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n    -+\t\t\t\texit 1\n    -+\t\t\tfi\n    -+\t\t) >server.log 2>&1 &\n    ++\t\t\"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n    ++\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n    ++\t\t\t\"$TRASH_DIRECTORY/first-ready\" >server.log 2>&1 &\n     +\t\tserver_pid=$!\n     +\t} &&\n     +\ttest_when_finished \"\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\t\tkill $second_pid 2>/dev/null || :\n     +\t\twait $second_pid 2>/dev/null || :\n     +\t\" &&\n    -+\twait \"$server_pid\" &&\n    -+\twait \"$first_pid\" &&\n     +\twait \"$second_pid\" &&\n    ++\twait \"$first_pid\" &&\n    ++\twait \"$server_pid\" &&\n     +\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n     +\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n     +\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n3:  1ee5d7e027 = 3:  d9063deb60 fetch-pack: accept \"pack\" output for packfile URIs\n\nbase-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200\n-- \n2.55.0.openai.131.g83a728de1eb6\n"},{"id":"548862","messageId":"a6a40b80461377452a0b2c9204c3a659ab60a7d5.1784874850.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784874850.git.tnyman@openai.com","subject":"[PATCH v4 1/3] http-fetch: correct --index-pack-arg documentation","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-24T08:14:23Z","receivedAt":"2026-07-24T08:14:25Z","isPatch":true,"body":"The --packfile mode accepts one --index-pack-arg=<arg> option per\nargument passed to index-pack, but its documentation and option\ndependency errors still refer to the plural --index-pack-args form.\n\nCorrect the spelling and describe the repeatable per-argument form.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc | 8 ++++----\n http-fetch.c                      | 4 ++--\n 2 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 2200f073c4..09b5d675ee 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -50,11 +50,11 @@ commit-id::\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n \tThe hash is used to determine the name of the temporary file and is\n \tarbitrary. The output of index-pack is printed to stdout. Requires\n-\t--index-pack-args.\n+\tone or more --index-pack-arg options.\n \n---index-pack-args=<args>::\n-\tFor internal use only. The command to run on the contents of the\n-\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n+--index-pack-arg=<arg>::\n+\tFor internal use only. An argument to the command run on the contents\n+\tof the downloaded pack. This option can be specified multiple times.\n \n --recover::\n \tVerify that everything reachable from target is fetched.  Used after\ndiff --git a/http-fetch.c b/http-fetch.c\nindex f9b6ecb061..601a77c3c1 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)\n \n \tif (packfile) {\n \t\tif (!index_pack_args.nr)\n-\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n+\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n \n \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n \t\t\t\t      index_pack_args.v);\n@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)\n \t}\n \n \tif (index_pack_args.nr)\n-\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n+\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n \n \tif (commits_on_stdin) {\n \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548864","messageId":"144c98cdfa4492206db0f1dd40cc43c47f376673.1784874850.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784874850.git.tnyman@openai.com","subject":"[PATCH v4 2/3] http: avoid concurrent appends to partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-24T08:14:24Z","receivedAt":"2026-07-24T08:14:27Z","isPatch":true,"body":"Pack requests stage downloads in a predictable partial-pack file so an\ninterrupted transfer can be resumed. Both packfile URI and ordinary dumb\nHTTP requests use this staging path. Opening it in append mode forces\neach write to the current end of the file, so concurrent responses can\nappend duplicate data and corrupt the pack.\n\nOpen the partial pack read-write without O_APPEND and seek once to its\ncurrent end. Each downloader then retains the offset matching the Range\nit requested. Because the staging key must uniquely identify immutable\npack contents, overlapping responses write the same bytes at the same\noffsets instead of extending the file with duplicate data.\n\nMinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\nexisting file. Create a missing partial pack exclusively, close it, and\nreopen it without O_CREAT so every retained descriptor permits another\ndownloader to unlink the staging path. Duplicate that descriptor for\nindex-pack instead of reopening the path after closing the stream;\nindex-pack installs its own pack and the shared staging file is only\nunlinked, never renamed. Accept HTTP 416 when a partial pack is already\ncomplete and let index-pack validate its contents.\n\nExercise resumed transfers, EOF ranges, overlapping 200 and 206\nresponses, and unlinking the staging path while index-pack still holds\nits descriptor. Clarify the staging-key documentation.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc |   5 +-\n http-fetch.c                      |   3 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 ++++---\n t/t5550-http-fetch-dumb.sh        | 250 ++++++++++++++++++++++++++++++\n 6 files changed, 295 insertions(+), 25 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 09b5d675ee..60ca91cf3a 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -48,8 +48,9 @@ commit-id::\n \tline (which is not expected in\n \tthis case), 'git http-fetch' fetches the packfile directly at the given\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n-\tThe hash is used to determine the name of the temporary file and is\n-\tarbitrary. The output of index-pack is printed to stdout. Requires\n+\tThe hash is used to determine the name of the temporary file. It need\n+\tnot be the pack hash, but it must uniquely identify the pack contents\n+\tfor resumption. The output of index-pack is printed to stdout. Requires\n \tone or more --index-pack-arg options.\n \n --index-pack-arg=<arg>::\ndiff --git a/http-fetch.c b/http-fetch.c\nindex 601a77c3c1..05f68f306a 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\tstruct url_info url;\n \t\t\tchar *nurl = url_normalize(preq->url, &url);\n \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\ndiff --git a/http-push.c b/http-push.c\nindex 60f6f8f054..ef8abe3908 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)\n \n \t} else if (request->state == RUN_FETCH_PACKED) {\n \t\tint fail = 1;\n-\t\tif (request->curl_result != CURLE_OK) {\n+\t\tif (request->curl_result != CURLE_OK &&\n+\t\t    request->http_code != 416) {\n \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n \t\t\t\trequest->url, curl_errorstr);\n \t\t} else {\ndiff --git a/http-walker.c b/http-walker.c\nindex b58a3b2a92..abafca84d6 100644\n--- a/http-walker.c\n+++ b/http-walker.c\n@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n \t\t\t      curl_errorstr);\n \t\t\tgoto abort;\ndiff --git a/http.c b/http.c\nindex caccf2108e..a0d399b274 100644\n--- a/http.c\n+++ b/http.c\n@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \tint tmpfile_fd;\n \tint ret = 0;\n \n+\t/* Another downloader may unlink the staging path while we index it. */\n+\ttmpfile_fd = xdup(fileno(preq->packfile));\n \tfclose(preq->packfile);\n \tpreq->packfile = NULL;\n-\n-\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n+\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n+\t\tdie_errno(\"unable to seek local file %s for pack\",\n+\t\t\t  preq->tmpfile.buf);\n \n \tip.git_cmd = 1;\n \tip.in = tmpfile_fd;\n@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \telse\n \t\tip.no_stdout = 1;\n \n-\tif (run_command(&ip)) {\n+\tif (run_command(&ip))\n \t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\n-cleanup:\n-\tclose(tmpfile_fd);\n \tunlink(preq->tmpfile.buf);\n \treturn ret;\n }\n@@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(\n struct http_pack_request *new_direct_http_pack_request(\n \tconst unsigned char *packed_git_hash, char *url)\n {\n-\toff_t prev_posn = 0;\n+\toff_t prev_posn;\n \tstruct http_pack_request *preq;\n+\tint fd;\n \n \tCALLOC_ARRAY(preq, 1);\n \tstrbuf_init(&preq->tmpfile, 0);\n-\n \tpreq->url = url;\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n-\tif (!preq->packfile) {\n-\t\terror(\"Unable to open local file %s for pack\",\n-\t\t      preq->tmpfile.buf);\n+\t/*\n+\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n+\t * existing file; reopen a newly created file so others may unlink it.\n+\t */\n+\tfor (;;) {\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n+\t\tif (fd >= 0 || errno != ENOENT)\n+\t\t\tbreak;\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n+\t\tif (fd >= 0) {\n+\t\t\tclose(fd);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (errno != EEXIST)\n+\t\t\tbreak;\n+\t}\n+\tif (fd < 0) {\n+\t\terror_errno(\"unable to open local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n \t\tgoto abort;\n \t}\n+\tprev_posn = lseek(fd, 0, SEEK_END);\n+\tif (prev_posn < 0) {\n+\t\terror_errno(\"unable to seek local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tclose(fd);\n+\t\tgoto abort;\n+\t}\n+\tpreq->packfile = xfdopen(fd, \"w\");\n \n \tpreq->slot = get_active_slot();\n \tpreq->headers = object_request_headers();\n@@ -2762,12 +2783,7 @@ struct http_pack_request *new_direct_http_pack_request(\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);\n \n-\t/*\n-\t * If there is data present from a previous transfer attempt,\n-\t * resume where it left off\n-\t */\n-\tprev_posn = ftello(preq->packfile);\n-\tif (prev_posn>0) {\n+\tif (prev_posn > 0) {\n \t\tif (http_is_verbose)\n \t\t\tfprintf(stderr,\n \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex f00eeae48f..dcb9667eeb 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -293,6 +293,256 @@ test_expect_success 'http-fetch --packfile' '\n \tgit -C packfileclient cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile resumes a partial download' '\n+\tgit init packfileclient-resume &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/resume.trace\" \\\n+\tgit -C packfileclient-resume http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=index-pack --index-pack-arg=--stdin \\\n+\t\t--index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_grep \"Range: bytes=64-\" resume.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success 'http-fetch --packfile permits unlink while indexing' '\n+\tgit init packfileclient-unlink &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\twrite_script git-unlink-index-pack <<-\\EOF &&\n+\ttest -f \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\trm \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\texec git index-pack \"$@\"\n+\tEOF\n+\ttest_when_finished \"rm -f git-unlink-index-pack\" &&\n+\tPATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\tGIT_TEST_PACK_TEMP=\"$TRASH_DIRECTORY/$tmpfile\" \\\n+\tgit -C packfileclient-unlink http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=unlink-index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-unlink cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n+\tgit init packfileclient-concurrent &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\tpackhash=$(basename \"$p\" .pack) &&\n+\tpackhash=${packhash#pack-} &&\n+\ttmpfile=\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tmkfifo first-ready first-continue &&\n+\texec 8<>first-ready &&\n+\texec 9<>first-continue &&\n+\twrite_script git-wait-index-pack <<-\\EOF &&\n+\techo ready >\"$GIT_TEST_WAIT_READY\" &&\n+\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n+\texec git index-pack \"$@\"\n+\tEOF\n+\t{\n+\t\t(\n+\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n+\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n+\t\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/first.trace\" \\\n+\t\t\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=wait-index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\techo continue >&9\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\texec 8>&-\n+\t\texec 9>&-\n+\t\trm -f first-ready first-continue git-wait-index-pack\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/second.trace\" \\\n+\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t--index-pack-arg=index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n+\techo continue >&9 &&\n+\twait \"$first_pid\" &&\n+\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect first.out &&\n+\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect second.out &&\n+\ttest_grep \"Range: bytes=64-\" first.trace &&\n+\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n+\ttest_grep \"416 Requested Range Not Satisfiable\" second.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n+\tgit init packfileclient-overlap &&\n+\tblob=$(test-tool genrandom pack-overlap 2m |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\thash-object -w --stdin) &&\n+\tpackhash=$(printf \"%s\\n\" \"$blob\" |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\tpack-objects \"$TRASH_DIRECTORY/overlap-pack\") &&\n+\tpack=\"$TRASH_DIRECTORY/overlap-pack-$packhash.pack\" &&\n+\ttmpfile=\"packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tmkfifo server-ready first-ready &&\n+\texec 7<>server-ready &&\n+\texec 8<>first-ready &&\n+\twrite_script slow-pack-server \"$PERL_PATH\" <<-\\EOF &&\n+\tuse strict;\n+\tuse warnings;\n+\tuse IO::Socket::INET;\n+\n+\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n+\tmy $completed = 0;\n+\tEND {\n+\t\tif (!$completed) {\n+\t\t\tsignal_ready($server_ready, \"failed\");\n+\t\t\tsignal_ready($first_ready, \"failed\");\n+\t\t}\n+\t}\n+\n+\t$SIG{ALRM} = sub { die \"timed out serving concurrent pack requests\\n\" };\n+\talarm 60;\n+\n+\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n+\tmy $pack = do { local $/; <$in> };\n+\tclose($in) or die \"close $packfile: $!\";\n+\tmy $server = IO::Socket::INET->new(LocalAddr => \"127.0.0.1\",\n+\t\tLocalPort => 0, Proto => \"tcp\", Listen => 2, ReuseAddr => 1)\n+\t\tor die \"listen: $!\";\n+\n+\tsub signal_ready {\n+\t\tmy ($file, $value) = @_;\n+\t\topen(my $out, \">\", $file) or die \"open $file: $!\";\n+\t\tprint $out \"$value\\n\" or die \"write $file: $!\";\n+\t\tclose($out) or die \"close $file: $!\";\n+\t}\n+\n+\tsub write_all {\n+\t\tmy ($out, $data) = @_;\n+\t\tmy $offset = 0;\n+\t\twhile ($offset < length($data)) {\n+\t\t\tmy $written = syswrite($out, $data,\n+\t\t\t\tlength($data) - $offset, $offset);\n+\t\t\tdefined($written) && $written or die \"write response: $!\";\n+\t\t\t$offset += $written;\n+\t\t}\n+\t}\n+\n+\tsub start_response {\n+\t\tmy $out = $server->accept() or die \"accept: $!\";\n+\t\t<$out> or die \"read request: $!\";\n+\t\tmy $start = 0;\n+\t\twhile (<$out>) {\n+\t\t\tlast if /^\\r?\\n$/;\n+\t\t\t$start = $1 if /^Range: bytes=(\\d+)-/i;\n+\t\t}\n+\t\t$start < length($pack) or die \"invalid range $start\";\n+\t\tmy $length = length($pack) - $start;\n+\t\tmy $middle = int($length / 2);\n+\t\tmy $status = $start ? \"206 Partial Content\" : \"200 OK\";\n+\t\tmy $headers = \"HTTP/1.1 $status\\r\\n\" .\n+\t\t\t\"Content-Length: $length\\r\\n\" .\n+\t\t\t($start ? \"Content-Range: bytes $start-\" .\n+\t\t\t\t(length($pack) - 1) . \"/\" . length($pack) . \"\\r\\n\" : \"\") .\n+\t\t\t\"Connection: close\\r\\n\\r\\n\";\n+\t\twrite_all($out, $headers);\n+\t\twrite_all($out, substr($pack, $start, $middle));\n+\t\treturn ($out, $start + $middle);\n+\t}\n+\n+\tsignal_ready($server_ready, $server->sockport());\n+\tmy ($first, $first_pos) = start_response();\n+\tsignal_ready($first_ready, \"ready\");\n+\tmy ($second, $second_pos) = start_response();\n+\twrite_all($first, substr($pack, $first_pos));\n+\twrite_all($second, substr($pack, $second_pos));\n+\tclose($first) or die \"close first response: $!\";\n+\tclose($second) or die \"close second response: $!\";\n+\t$completed = 1;\n+\talarm 0;\n+\tEOF\n+\t{\n+\t\t\"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n+\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n+\t\t\t\"$TRASH_DIRECTORY/first-ready\" >server.log 2>&1 &\n+\t\tserver_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $server_pid 2>/dev/null || :\n+\t\twait $server_pid 2>/dev/null || :\n+\t\texec 7>&-\n+\t\texec 8>&-\n+\t\trm -f server-ready first-ready slow-pack-server\n+\t\" &&\n+\tread port <&7 &&\n+\turl=\"http://127.0.0.1:$port/pack\" &&\n+\t{\n+\t\t(\n+\t\t\tif ! GIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-first.trace\" \\\n+\t\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$url\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\ttest_path_is_file \"$tmpfile\" &&\n+\ttest -s \"$tmpfile\" &&\n+\t{\n+\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n+\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\"$url\" >second.out &\n+\t\tsecond_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $second_pid 2>/dev/null || :\n+\t\twait $second_pid 2>/dev/null || :\n+\t\" &&\n+\twait \"$second_pid\" &&\n+\twait \"$first_pid\" &&\n+\twait \"$server_pid\" &&\n+\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n+\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n+\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n+\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n+\tsort first.out second.out >actual &&\n+\ttest_cmp expect actual &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-overlap cat-file -e \"$blob\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548863","messageId":"d9063deb60354eb731e34c453cd6730e1098f905.1784874850.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784874850.git.tnyman@openai.com","subject":"[PATCH v4 3/3] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-24T08:14:25Z","receivedAt":"2026-07-24T08:14:28Z","isPatch":true,"body":"When index-pack finds an existing keep file it reports pack rather than\nkeep. Accept either result from http-fetch, and only register a keep\nlockfile when this fetch created it.\n\nRead the pack/keep prefix and hash without consuming any following fsck\noutput, validate the reported pack hash against the advertised hash, and\nexercise a packfile URI fetch with a pre-existing keep file.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n fetch-pack.c           | 33 ++++++++++++++++++---------------\n t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 15 deletions(-)\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 29c41132ee..e9f24fbd63 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t}\n \n \tfor (i = 0; i < packfile_uris.nr; i++) {\n+\t\tbool created_keep;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 1];\n+\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n-\t\t    memcmp(packname, \"keep\\t\", 5))\n-\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n+\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n+\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n+\t\t     memcmp(packhash, \"pack\\t\", 5)))\n+\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n+\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n \n-\t\tif (read_in_full(cmd.out, packname,\n-\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n-\t\t    packname[the_hash_algo->hexsz] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n-\n-\t\tpackname[the_hash_algo->hexsz] = '\\0';\n+\t\tif (read_in_full(cmd.out, packhash,\n+\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n+\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n+\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n+\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n \n \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n \n@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (finish_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to finish http-fetch\");\n \n-\t\tif (memcmp(packfile_uris.items[i].string, packname,\n+\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n \t\t\t   the_hash_algo->hexsz))\n \t\t\tdie(\"fetch-pack: pack downloaded from %s does not match expected hash %.*s\",\n \t\t\t    uri, (int) the_hash_algo->hexsz,\n \t\t\t    packfile_uris.items[i].string);\n \n-\t\tstring_list_append_nodup(pack_lockfiles,\n-\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n-\t\t\t\t\t\t repo_get_object_directory(the_repository),\n-\t\t\t\t\t\t packname));\n+\t\tif (created_keep)\n+\t\t\tstring_list_append_nodup(pack_lockfiles,\n+\t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n+\t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n+\t\t\t\t\t\t\t packhash));\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\ndiff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh\nindex 74a2b7730b..0f05286de8 100755\n--- a/t/t5702-protocol-v2.sh\n+++ b/t/t5702-protocol-v2.sh\n@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '\n \t\tfetch \"$HTTPD_URL/smart/http_parent\"\n '\n \n+test_expect_success 'packfile URI preserves an existing keep file' '\n+\tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n+\trm -rf \"$P\" http_child keep.expect &&\n+\n+\tgit init \"$P\" &&\n+\tgit -C \"$P\" config uploadpack.allowsidebandall true &&\n+\n+\techo my-blob >\"$P/my-blob\" &&\n+\tgit -C \"$P\" add my-blob &&\n+\tgit -C \"$P\" commit -m x &&\n+\tconfigure_exclusion \"$P\" my-blob >h &&\n+\n+\tgit init http_child &&\n+\tpackhash=$(cat packh) &&\n+\tkeep=\"http_child/.git/objects/pack/pack-$packhash.keep\" &&\n+\techo pre-existing >\"$keep\" &&\n+\tcp \"$keep\" keep.expect &&\n+\n+\tGIT_TEST_SIDEBAND_ALL=1 \\\n+\tgit -C http_child -c protocol.version=2 \\\n+\t\t-c fetch.uriprotocols=http,https \\\n+\t\tfetch \"$HTTPD_URL/smart/http_parent\" &&\n+\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.pack\" &&\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.idx\" &&\n+\ttest_cmp keep.expect \"$keep\" &&\n+\tgit -C http_child cat-file -e \"$(cat h)\"\n+'\n+\n test_expect_success 'fetching with valid packfile URI but invalid hash fails' '\n \tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n \trm -rf \"$P\" http_child log &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548929","messageId":"amPbRCAOLr-pSWfj@com-79390","threadId":"65988","inReplyTo":"a6a40b80461377452a0b2c9204c3a659ab60a7d5.1784874850.git.tnyman@openai.com","subject":"Re: [PATCH v4 1/3] http-fetch: correct --index-pack-arg documentation","fromName":"Taylor Blau","fromEmail":"ttaylorr@openai.com","sentAt":"2026-07-24T21:38:12Z","receivedAt":"2026-07-24T21:38:16Z","isPatch":true,"body":"On Fri, Jul 24, 2026 at 01:14:23AM -0700, Ted Nyman wrote:\n> The --packfile mode accepts one --index-pack-arg=<arg> option per\n> argument passed to index-pack, but its documentation and option\n> dependency errors still refer to the plural --index-pack-args form.\n\nGood find, it looks like this dates all the way back to 27e35ba6c6\n(http-fetch: allow custom index-pack args, 2021-02-22). Thanks for\ntaking the time to correct it.\n\n> diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\n> index 2200f073c4..09b5d675ee 100644\n> --- a/Documentation/git-http-fetch.adoc\n> +++ b/Documentation/git-http-fetch.adoc\n> @@ -50,11 +50,11 @@ commit-id::\n>  \tURL and uses index-pack to generate corresponding .idx and .keep files.\n>  \tThe hash is used to determine the name of the temporary file and is\n>  \tarbitrary. The output of index-pack is printed to stdout. Requires\n> -\t--index-pack-args.\n> +\tone or more --index-pack-arg options.\n>\n> ---index-pack-args=<args>::\n> -\tFor internal use only. The command to run on the contents of the\n> -\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n> +--index-pack-arg=<arg>::\n> +\tFor internal use only. An argument to the command run on the contents\n> +\tof the downloaded pack. This option can be specified multiple times.\n\nInteresting. The plural \"--index-pack-args\" form says that it specifies\nthe command to run on the downloaded pack, as well as arguments which\nare separated by spaces. Two thoughts:\n\n - I think the \"arguments are URL-encoded separated by spaces\" claim was\n   not true even in 27e35ba6c6, so dropping that seems like a strict\n   improvement to me.\n\n - The new form says \"An argument to the command run on [...]\", but I\n   believe that this option is also used to specify the name of the\n   command to run itself. I wonder if it may be worth saying something\n   like \"The first instance specifies the command to run. Subsequent\n   occurrences specify its arguments.\"\n\n> diff --git a/http-fetch.c b/http-fetch.c\n> index f9b6ecb061..601a77c3c1 100644\n> --- a/http-fetch.c\n> +++ b/http-fetch.c\n\nChanges in this file look reasonable. Likewise, it makes sense that we\ndo not have any changes in the test suite, since this option did not\nexist in a plural in the first place ;-).\n\nThanks,\nTaylor\n"},{"id":"548930","messageId":"amPdMQH3QRLnDpl0@com-79390","threadId":"65988","inReplyTo":"d9063deb60354eb731e34c453cd6730e1098f905.1784874850.git.tnyman@openai.com","subject":"Re: [PATCH v4 3/3] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Taylor Blau","fromEmail":"ttaylorr@openai.com","sentAt":"2026-07-24T21:46:25Z","receivedAt":"2026-07-24T21:46:29Z","isPatch":true,"body":"On Fri, Jul 24, 2026 at 01:14:25AM -0700, Ted Nyman wrote:\n> When index-pack finds an existing keep file it reports pack rather than\n> keep. Accept either result from http-fetch, and only register a keep\n> lockfile when this fetch created it.\n>\n> Read the pack/keep prefix and hash without consuming any following fsck\n> output, validate the reported pack hash against the advertised hash, and\n> exercise a packfile URI fetch with a pre-existing keep file.\n>\n> Signed-off-by: Ted Nyman <tnyman@openai.com>\n> ---\n>  fetch-pack.c           | 33 ++++++++++++++++++---------------\n>  t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n>  2 files changed, 49 insertions(+), 15 deletions(-)\n>\n> diff --git a/fetch-pack.c b/fetch-pack.c\n> index 29c41132ee..e9f24fbd63 100644\n> --- a/fetch-pack.c\n> +++ b/fetch-pack.c\n> @@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n>  \t}\n>\n>  \tfor (i = 0; i < packfile_uris.nr; i++) {\n> +\t\tbool created_keep;\n>  \t\tint j;\n>  \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n> -\t\tchar packname[GIT_MAX_HEXSZ + 1];\n> +\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n\nOK, so we keep track of whether or not we got \"keep\" as part of the\noutput.\n\nWhile here, \"packname\" is renamed to \"packhash\", which I think is\nreasonable, especially to indicate that the buffer is sized accordingly.\nWe happen to read the preceding \"pack\" or \"keep\" into that same buffer,\nwhich I think is fine. If we wanted to be pedantic we could read that\ninto a separate buffer, but I don't think such separation is necessary.\n\n>  \t\tconst char *uri = packfile_uris.items[i].string +\n>  \t\t\tthe_hash_algo->hexsz + 1;\n>\n> @@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n>  \t\tif (start_command(&cmd))\n>  \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n>\n> -\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n> -\t\t    memcmp(packname, \"keep\\t\", 5))\n> -\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n> +\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n> +\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n> +\t\t     memcmp(packhash, \"pack\\t\", 5)))\n> +\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n> +\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n\nMakes sense.\n\n>\n> -\t\tif (read_in_full(cmd.out, packname,\n> -\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n> -\t\t    packname[the_hash_algo->hexsz] != '\\n')\n> -\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n> -\n> -\t\tpackname[the_hash_algo->hexsz] = '\\0';\n> +\t\tif (read_in_full(cmd.out, packhash,\n> +\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n> +\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n> +\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n> +\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n\nLikewise. The rest of this file and the test also look good.\n\nThanks,\nTaylor\n"},{"id":"548936","messageId":"20260725090910.GA1438796@coredump.intra.peff.net","threadId":"65988","inReplyTo":"xmqqldb19evx.fsf@gitster.g","subject":"Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-25T09:09:10Z","receivedAt":"2026-07-25T09:09:12Z","isPatch":true,"body":"On Thu, Jul 23, 2026 at 09:43:14PM -0700, Junio C Hamano wrote:\n\n> When merged into 'seen', this topic seems to cause t5550 to hang\n> fairly consistently.  It is not surprising, considering that the\n> topic adds roughly 240 lines to the test script in question.  It is\n> entirely possible that we are seeing an existing breakage from\n> another topic in 'seen' that is exposed by the additional tests.\n\nI didn't get any hang locally, but running t5550 with --stress causes\naround half of the runs to fail immediately. That continues to be true\nwith v4. So there is presumably some race condition still present.\n\nThe failing test is the big one (34) and the failing command is the\n\"test -s $tmpfile\" call. It looks like the pack has already been indexed\n(at least by the time I look at the on-disk state of a failed example).\n\nI don't immediately see the issue, though. I could believe that extra\nload fakes out any sleep-based timing tricks, but it looks like the test\ntries to use FIFOs to do everything deterministically.\n\nDiffing the overlap-first.trace file between a working case and a\nfailing one, I see (skipping past uninteresting port differences) this\nhunk at the end:\n\n  @@ -17,4 +17,5 @@\n   <= Recv header: Connection: close\n   <= Recv header, 0000000002 bytes (0x00000002)\n   <= Recv header:\n  -== Info: shutting down connection #0\n  +== Info: end of response with 1048917 bytes missing\n  +== Info: closing connection #0\n\nSo curl sees a hangup on the first connection (even though the second\none hasn't even started yet!). I'm not sure why, though. There's nothing\nuseful in the server.log file. I tried stracing the server process but\nit didn't show much of interest. Both cases write \"ready\" to\nfirst-ready, and then the success case immediately sees an accept() for\nthe second connection. The failing case waits in accept() and then\neventually calls SIGALRM (which is way after the failure happens; the\ntest has already bailed and so the second connection never comes in).\n\nSo from the perspective of the server process, everything is fine, but\ncurl complains that it didn't get all of the bytes. Weird. The strace\nshows both writing the first 1MB as expected. It's like the connection\ngets hung up for some reason, but I can't tell why or by whom.\n\n-Peff\n"},{"id":"548937","messageId":"20260725092154.GA1925154@coredump.intra.peff.net","threadId":"65988","inReplyTo":"20260725090910.GA1438796@coredump.intra.peff.net","subject":"Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-25T09:21:54Z","receivedAt":"2026-07-25T09:21:56Z","isPatch":true,"body":"On Sat, Jul 25, 2026 at 05:09:11AM -0400, Jeff King wrote:\n\n> So from the perspective of the server process, everything is fine, but\n> curl complains that it didn't get all of the bytes. Weird. The strace\n> shows both writing the first 1MB as expected. It's like the connection\n> gets hung up for some reason, but I can't tell why or by whom.\n\nHmph. I tried stracing on the client side, and we indeed see an EOF\non the socket:\n\n  recvfrom(12, \"\", 16384, 0, NULL, NULL) = 0\n\nI'm really puzzled why that is the case, though, as the server side did\nnot close() or exit.\n\nI do wonder if it would be possible to write this script so that it\ntriggers via apache, like the rest of our http tests. Doing our own\nsocket handling here is a potential source of bugs.\n\n-Peff\n"},{"id":"548938","messageId":"20260725100251.GA1933232@coredump.intra.peff.net","threadId":"65988","inReplyTo":"20260725092154.GA1925154@coredump.intra.peff.net","subject":"Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-25T10:02:51Z","receivedAt":"2026-07-25T10:02:53Z","isPatch":true,"body":"On Sat, Jul 25, 2026 at 05:21:54AM -0400, Jeff King wrote:\n\n> On Sat, Jul 25, 2026 at 05:09:11AM -0400, Jeff King wrote:\n> \n> > So from the perspective of the server process, everything is fine, but\n> > curl complains that it didn't get all of the bytes. Weird. The strace\n> > shows both writing the first 1MB as expected. It's like the connection\n> > gets hung up for some reason, but I can't tell why or by whom.\n> \n> Hmph. I tried stracing on the client side, and we indeed see an EOF\n> on the socket:\n> \n>   recvfrom(12, \"\", 16384, 0, NULL, NULL) = 0\n> \n> I'm really puzzled why that is the case, though, as the server side did\n> not close() or exit.\n\nSo I'm still puzzled by all of this, but I think it's mostly a red\nherring with respect to the actual \"test -s\" race.\n\nThe server tells us when it has written the first 1MB to the client, and\nthen we check that \"test -s\" is showing something in the on-disk\ntempfile we're downloading.  But there's no guarantee that just because\nthe server called write() that client has yet received the data, let\nalone written it to disk.\n\nI can't think of a synchronization point we could use here. We're\nwaiting on curl to have passed the bytes to fwrite() and for it to have\nactually synced to disk. We either have to poll or modify http.c to\nwrite \"yes, we got some bytes!\" to a fifo. Both are pretty gross.\n\nI wonder if we could just drop that \"test -s\" entirely. We'd _usually_\nsee some bytes written before the second request starts. But it's OK if\nwe don't. It just means the test is working in the reverse order (the\nsecond request may write its bytes first, and then the first one is the\none \"overwriting\" it). I.e., the two are symmetric from our perspective.\n\n-Peff\n"},{"id":"548939","messageId":"20260725101043.GA2171844@coredump.intra.peff.net","threadId":"65988","inReplyTo":"20260725100251.GA1933232@coredump.intra.peff.net","subject":"Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-25T10:10:43Z","receivedAt":"2026-07-25T10:10:45Z","isPatch":true,"body":"On Sat, Jul 25, 2026 at 06:02:51AM -0400, Jeff King wrote:\n\n> I wonder if we could just drop that \"test -s\" entirely. We'd _usually_\n> see some bytes written before the second request starts. But it's OK if\n> we don't. It just means the test is working in the reverse order (the\n> second request may write its bytes first, and then the first one is the\n> one \"overwriting\" it). I.e., the two are symmetric from our perspective.\n\nYeah, doing this:\n\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex dcb9667eeb..07aa218049 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -516,7 +516,6 @@ test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt a\n \tread ready <&8 &&\n \ttest \"$ready\" = ready &&\n \ttest_path_is_file \"$tmpfile\" &&\n-\ttest -s \"$tmpfile\" &&\n \t{\n \t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n \t\tGIT_TRACE_CURL_NO_DATA=1 \\\n@@ -533,9 +532,6 @@ test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt a\n \twait \"$second_pid\" &&\n \twait \"$first_pid\" &&\n \twait \"$server_pid\" &&\n-\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n-\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n-\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n \tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n \tsort first.out second.out >actual &&\n \ttest_cmp expect actual &&\n\nis enough to make it pass reliably under --stress for me. We have to\ndrop the trace greps, because we don't actually know whether each\nrequest will use a range or not. We'd _usually_ see a range for the\nsecond one, but it's possible it might still see a zero-byte file. I\nguess we probably see a \"200\" reliably for the first request, but it's\nnot all that interesting.\n\nWe can leave the test_path_is_file check, because we open the file\nbefore making the request (it is only the actual writing of bytes that\nis racy).\n\n-Peff\n"},{"id":"548970","messageId":"xmqqik6311oj.fsf@gitster.g","threadId":"65988","inReplyTo":"20260725100251.GA1933232@coredump.intra.peff.net","subject":"Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-25T16:20:12Z","receivedAt":"2026-07-25T16:20:15Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> I can't think of a synchronization point we could use here. We're\n> waiting on curl to have passed the bytes to fwrite() and for it to have\n> actually synced to disk. We either have to poll or modify http.c to\n> write \"yes, we got some bytes!\" to a fifo. Both are pretty gross.\n>\n> I wonder if we could just drop that \"test -s\" entirely. We'd _usually_\n> see some bytes written before the second request starts. But it's OK if\n> we don't. It just means the test is working in the reverse order (the\n> second request may write its bytes first, and then the first one is the\n> one \"overwriting\" it). I.e., the two are symmetric from our perspective.\n\nYeah, that sounds quite sensible.\nThanks for digging.\n"},{"id":"548980","messageId":"cover.1785047139.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1784874850.git.tnyman@openai.com","subject":"[PATCH v5 0/3] packfile URIs: support concurrent downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-26T06:44:45Z","receivedAt":"2026-07-26T06:44:50Z","isPatch":true,"body":"Packfile URI and dumb HTTP downloads stage packs at\nobjects/pack/pack-<hash>.pack.temp so an interrupted transfer can\nresume. Opening that file in append mode forces every write to its\ncurrent end. Two Git processes fetching the same pack into one object\ndatabase can therefore append duplicate data and corrupt the pack.\n\nThe first patch separates the unrelated --index-pack-arg documentation\nand error-message correction requested during review.\n\nThe second patch keeps the predictable staging name but removes append\nmode. Each downloader seeks once to the current end, requests the\ncorresponding Range, and writes using its own descriptor offset. Since\nthe staging key must identify immutable pack contents, overlapping\nresponses write identical bytes at identical offsets. There is no need\nfor pwrite(2) or cross-process coordination, and resumption continues to\nwork for both packfile URI and ordinary dumb HTTP downloads.\n\nA downloader can also find that the partial pack has completed and\nrequest a range starting at EOF. Servers may respond with HTTP 416 in\nthat case. Treat the response as a completed download and let\nindex-pack validate the pack.\n\nOn MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an\nexisting file. Create a missing staging file exclusively, close it, and\nreopen it without O_CREAT so every retained descriptor permits another\ndownloader to unlink the path. Keep the open descriptor for index-pack;\nit installs its own pack, so the shared staging file is only unlinked,\nnever renamed.\n\nThe third patch handles the related .keep race. When another process has\nalready created the keep file, index-pack reports \"pack<TAB><hash>\"\ninstead of \"keep<TAB><hash>\". Accept both successful forms and remove\nonly keep files created by the current process. Read only the prefix and\nhash so any following fsck output remains available to fetch-pack.\n\nThe tests cover resumption, a completed partial returning 416,\noverlapping downloads, unlinking the staging path while index-pack holds\nits descriptor, and a pre-existing .keep file. The unlink test does not\nrequire FIFOs, so it can exercise MinGW's sharing behavior even though\nthe concurrent-download tests are skipped there.\n\nChanges since v4:\n\n  * Clarify that the first --index-pack-arg specifies the command and\n    subsequent instances specify its arguments.\n  * Drop assumptions about which concurrent response reaches the\n    staging file first. Either write order exercises the same\n    overlapping-download behavior.\n  * No production code changes.\n\nThe overlapping-download test passes 240 runs with 12 parallel stress\njobs.\n\nThe v4 discussion is at:\n\n  https://lore.kernel.org/git/cover.1784874850.git.tnyman@openai.com/\n\nTed Nyman (3):\n  http-fetch: correct --index-pack-arg documentation\n  http: avoid concurrent appends to partial packs\n  fetch-pack: accept \"pack\" output for packfile URIs\n\n Documentation/git-http-fetch.adoc |  14 +-\n fetch-pack.c                      |  33 ++--\n http-fetch.c                      |   7 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 ++++---\n t/t5550-http-fetch-dumb.sh        | 246 ++++++++++++++++++++++++++++++\n t/t5702-protocol-v2.sh            |  31 ++++\n 8 files changed, 347 insertions(+), 46 deletions(-)\n\nRange-diff against v4:\n1:  a6a40b8046 ! 1:  a79af009ea http-fetch: correct --index-pack-arg documentation\n    @@ Documentation/git-http-fetch.adoc: commit-id::\n     -\tFor internal use only. The command to run on the contents of the\n     -\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n     +--index-pack-arg=<arg>::\n    -+\tFor internal use only. An argument to the command run on the contents\n    -+\tof the downloaded pack. This option can be specified multiple times.\n    ++\tFor internal use only. The first instance specifies the command run on\n    ++\tthe contents of the downloaded pack. Subsequent instances specify its\n    ++\targuments.\n      \n      --recover::\n      \tVerify that everything reachable from target is fetched.  Used after\n2:  144c98cdfa ! 2:  d9667c93b0 http: avoid concurrent appends to partial packs\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\tread ready <&8 &&\n     +\ttest \"$ready\" = ready &&\n     +\ttest_path_is_file \"$tmpfile\" &&\n    -+\ttest -s \"$tmpfile\" &&\n     +\t{\n     +\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n     +\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\twait \"$second_pid\" &&\n     +\twait \"$first_pid\" &&\n     +\twait \"$server_pid\" &&\n    -+\ttest_grep \"HTTP/[0-9.]* 200\" overlap-first.trace &&\n    -+\ttest_grep \"Range: bytes=[1-9][0-9]*-\" overlap-second.trace &&\n    -+\ttest_grep \"HTTP/[0-9.]* 206\" overlap-second.trace &&\n     +\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n     +\tsort first.out second.out >actual &&\n     +\ttest_cmp expect actual &&\n3:  d9063deb60 = 3:  fee6f292cb fetch-pack: accept \"pack\" output for packfile URIs\n\nbase-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200\n-- \n2.55.0.openai.131.g83a728de1eb6\n"},{"id":"548981","messageId":"a79af009eaedf607ce5110d7b0f880af0a2764ba.1785047139.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785047139.git.tnyman@openai.com","subject":"[PATCH v5 1/3] http-fetch: correct --index-pack-arg documentation","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-26T06:44:46Z","receivedAt":"2026-07-26T06:44:51Z","isPatch":true,"body":"The --packfile mode accepts one --index-pack-arg=<arg> option per\nargument passed to index-pack, but its documentation and option\ndependency errors still refer to the plural --index-pack-args form.\n\nCorrect the spelling and describe the repeatable per-argument form.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc | 9 +++++----\n http-fetch.c                      | 4 ++--\n 2 files changed, 7 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 2200f073c4..12036e65e9 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -50,11 +50,12 @@ commit-id::\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n \tThe hash is used to determine the name of the temporary file and is\n \tarbitrary. The output of index-pack is printed to stdout. Requires\n-\t--index-pack-args.\n+\tone or more --index-pack-arg options.\n \n---index-pack-args=<args>::\n-\tFor internal use only. The command to run on the contents of the\n-\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n+--index-pack-arg=<arg>::\n+\tFor internal use only. The first instance specifies the command run on\n+\tthe contents of the downloaded pack. Subsequent instances specify its\n+\targuments.\n \n --recover::\n \tVerify that everything reachable from target is fetched.  Used after\ndiff --git a/http-fetch.c b/http-fetch.c\nindex f9b6ecb061..601a77c3c1 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)\n \n \tif (packfile) {\n \t\tif (!index_pack_args.nr)\n-\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n+\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n \n \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n \t\t\t\t      index_pack_args.v);\n@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)\n \t}\n \n \tif (index_pack_args.nr)\n-\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n+\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n \n \tif (commits_on_stdin) {\n \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548982","messageId":"d9667c93b03d1a71df55a33f90538b31afd08677.1785047139.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785047139.git.tnyman@openai.com","subject":"[PATCH v5 2/3] http: avoid concurrent appends to partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-26T06:44:47Z","receivedAt":"2026-07-26T06:44:53Z","isPatch":true,"body":"Pack requests stage downloads in a predictable partial-pack file so an\ninterrupted transfer can be resumed. Both packfile URI and ordinary dumb\nHTTP requests use this staging path. Opening it in append mode forces\neach write to the current end of the file, so concurrent responses can\nappend duplicate data and corrupt the pack.\n\nOpen the partial pack read-write without O_APPEND and seek once to its\ncurrent end. Each downloader then retains the offset matching the Range\nit requested. Because the staging key must uniquely identify immutable\npack contents, overlapping responses write the same bytes at the same\noffsets instead of extending the file with duplicate data.\n\nMinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\nexisting file. Create a missing partial pack exclusively, close it, and\nreopen it without O_CREAT so every retained descriptor permits another\ndownloader to unlink the staging path. Duplicate that descriptor for\nindex-pack instead of reopening the path after closing the stream;\nindex-pack installs its own pack and the shared staging file is only\nunlinked, never renamed. Accept HTTP 416 when a partial pack is already\ncomplete and let index-pack validate its contents.\n\nExercise resumed transfers, EOF ranges, overlapping 200 and 206\nresponses, and unlinking the staging path while index-pack still holds\nits descriptor. Clarify the staging-key documentation.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc |   5 +-\n http-fetch.c                      |   3 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 ++++---\n t/t5550-http-fetch-dumb.sh        | 246 ++++++++++++++++++++++++++++++\n 6 files changed, 291 insertions(+), 25 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 12036e65e9..45e0d3d07c 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -48,8 +48,9 @@ commit-id::\n \tline (which is not expected in\n \tthis case), 'git http-fetch' fetches the packfile directly at the given\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n-\tThe hash is used to determine the name of the temporary file and is\n-\tarbitrary. The output of index-pack is printed to stdout. Requires\n+\tThe hash is used to determine the name of the temporary file. It need\n+\tnot be the pack hash, but it must uniquely identify the pack contents\n+\tfor resumption. The output of index-pack is printed to stdout. Requires\n \tone or more --index-pack-arg options.\n \n --index-pack-arg=<arg>::\ndiff --git a/http-fetch.c b/http-fetch.c\nindex 601a77c3c1..05f68f306a 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\tstruct url_info url;\n \t\t\tchar *nurl = url_normalize(preq->url, &url);\n \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\ndiff --git a/http-push.c b/http-push.c\nindex 60f6f8f054..ef8abe3908 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)\n \n \t} else if (request->state == RUN_FETCH_PACKED) {\n \t\tint fail = 1;\n-\t\tif (request->curl_result != CURLE_OK) {\n+\t\tif (request->curl_result != CURLE_OK &&\n+\t\t    request->http_code != 416) {\n \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n \t\t\t\trequest->url, curl_errorstr);\n \t\t} else {\ndiff --git a/http-walker.c b/http-walker.c\nindex b58a3b2a92..abafca84d6 100644\n--- a/http-walker.c\n+++ b/http-walker.c\n@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n \t\t\t      curl_errorstr);\n \t\t\tgoto abort;\ndiff --git a/http.c b/http.c\nindex caccf2108e..a0d399b274 100644\n--- a/http.c\n+++ b/http.c\n@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \tint tmpfile_fd;\n \tint ret = 0;\n \n+\t/* Another downloader may unlink the staging path while we index it. */\n+\ttmpfile_fd = xdup(fileno(preq->packfile));\n \tfclose(preq->packfile);\n \tpreq->packfile = NULL;\n-\n-\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n+\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n+\t\tdie_errno(\"unable to seek local file %s for pack\",\n+\t\t\t  preq->tmpfile.buf);\n \n \tip.git_cmd = 1;\n \tip.in = tmpfile_fd;\n@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \telse\n \t\tip.no_stdout = 1;\n \n-\tif (run_command(&ip)) {\n+\tif (run_command(&ip))\n \t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\n-cleanup:\n-\tclose(tmpfile_fd);\n \tunlink(preq->tmpfile.buf);\n \treturn ret;\n }\n@@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(\n struct http_pack_request *new_direct_http_pack_request(\n \tconst unsigned char *packed_git_hash, char *url)\n {\n-\toff_t prev_posn = 0;\n+\toff_t prev_posn;\n \tstruct http_pack_request *preq;\n+\tint fd;\n \n \tCALLOC_ARRAY(preq, 1);\n \tstrbuf_init(&preq->tmpfile, 0);\n-\n \tpreq->url = url;\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n-\tif (!preq->packfile) {\n-\t\terror(\"Unable to open local file %s for pack\",\n-\t\t      preq->tmpfile.buf);\n+\t/*\n+\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n+\t * existing file; reopen a newly created file so others may unlink it.\n+\t */\n+\tfor (;;) {\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n+\t\tif (fd >= 0 || errno != ENOENT)\n+\t\t\tbreak;\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n+\t\tif (fd >= 0) {\n+\t\t\tclose(fd);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (errno != EEXIST)\n+\t\t\tbreak;\n+\t}\n+\tif (fd < 0) {\n+\t\terror_errno(\"unable to open local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n \t\tgoto abort;\n \t}\n+\tprev_posn = lseek(fd, 0, SEEK_END);\n+\tif (prev_posn < 0) {\n+\t\terror_errno(\"unable to seek local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tclose(fd);\n+\t\tgoto abort;\n+\t}\n+\tpreq->packfile = xfdopen(fd, \"w\");\n \n \tpreq->slot = get_active_slot();\n \tpreq->headers = object_request_headers();\n@@ -2762,12 +2783,7 @@ struct http_pack_request *new_direct_http_pack_request(\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);\n \n-\t/*\n-\t * If there is data present from a previous transfer attempt,\n-\t * resume where it left off\n-\t */\n-\tprev_posn = ftello(preq->packfile);\n-\tif (prev_posn>0) {\n+\tif (prev_posn > 0) {\n \t\tif (http_is_verbose)\n \t\t\tfprintf(stderr,\n \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex f00eeae48f..07aa218049 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -293,6 +293,252 @@ test_expect_success 'http-fetch --packfile' '\n \tgit -C packfileclient cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile resumes a partial download' '\n+\tgit init packfileclient-resume &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/resume.trace\" \\\n+\tgit -C packfileclient-resume http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=index-pack --index-pack-arg=--stdin \\\n+\t\t--index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_grep \"Range: bytes=64-\" resume.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success 'http-fetch --packfile permits unlink while indexing' '\n+\tgit init packfileclient-unlink &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\twrite_script git-unlink-index-pack <<-\\EOF &&\n+\ttest -f \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\trm \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\texec git index-pack \"$@\"\n+\tEOF\n+\ttest_when_finished \"rm -f git-unlink-index-pack\" &&\n+\tPATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\tGIT_TEST_PACK_TEMP=\"$TRASH_DIRECTORY/$tmpfile\" \\\n+\tgit -C packfileclient-unlink http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=unlink-index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-unlink cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n+\tgit init packfileclient-concurrent &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\tpackhash=$(basename \"$p\" .pack) &&\n+\tpackhash=${packhash#pack-} &&\n+\ttmpfile=\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tmkfifo first-ready first-continue &&\n+\texec 8<>first-ready &&\n+\texec 9<>first-continue &&\n+\twrite_script git-wait-index-pack <<-\\EOF &&\n+\techo ready >\"$GIT_TEST_WAIT_READY\" &&\n+\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n+\texec git index-pack \"$@\"\n+\tEOF\n+\t{\n+\t\t(\n+\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n+\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n+\t\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/first.trace\" \\\n+\t\t\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=wait-index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\techo continue >&9\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\texec 8>&-\n+\t\texec 9>&-\n+\t\trm -f first-ready first-continue git-wait-index-pack\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/second.trace\" \\\n+\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n+\t\t--index-pack-arg=index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n+\techo continue >&9 &&\n+\twait \"$first_pid\" &&\n+\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect first.out &&\n+\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n+\ttest_cmp expect second.out &&\n+\ttest_grep \"Range: bytes=64-\" first.trace &&\n+\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n+\ttest_grep \"416 Requested Range Not Satisfiable\" second.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n+\tgit init packfileclient-overlap &&\n+\tblob=$(test-tool genrandom pack-overlap 2m |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\thash-object -w --stdin) &&\n+\tpackhash=$(printf \"%s\\n\" \"$blob\" |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\tpack-objects \"$TRASH_DIRECTORY/overlap-pack\") &&\n+\tpack=\"$TRASH_DIRECTORY/overlap-pack-$packhash.pack\" &&\n+\ttmpfile=\"packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tmkfifo server-ready first-ready &&\n+\texec 7<>server-ready &&\n+\texec 8<>first-ready &&\n+\twrite_script slow-pack-server \"$PERL_PATH\" <<-\\EOF &&\n+\tuse strict;\n+\tuse warnings;\n+\tuse IO::Socket::INET;\n+\n+\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n+\tmy $completed = 0;\n+\tEND {\n+\t\tif (!$completed) {\n+\t\t\tsignal_ready($server_ready, \"failed\");\n+\t\t\tsignal_ready($first_ready, \"failed\");\n+\t\t}\n+\t}\n+\n+\t$SIG{ALRM} = sub { die \"timed out serving concurrent pack requests\\n\" };\n+\talarm 60;\n+\n+\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n+\tmy $pack = do { local $/; <$in> };\n+\tclose($in) or die \"close $packfile: $!\";\n+\tmy $server = IO::Socket::INET->new(LocalAddr => \"127.0.0.1\",\n+\t\tLocalPort => 0, Proto => \"tcp\", Listen => 2, ReuseAddr => 1)\n+\t\tor die \"listen: $!\";\n+\n+\tsub signal_ready {\n+\t\tmy ($file, $value) = @_;\n+\t\topen(my $out, \">\", $file) or die \"open $file: $!\";\n+\t\tprint $out \"$value\\n\" or die \"write $file: $!\";\n+\t\tclose($out) or die \"close $file: $!\";\n+\t}\n+\n+\tsub write_all {\n+\t\tmy ($out, $data) = @_;\n+\t\tmy $offset = 0;\n+\t\twhile ($offset < length($data)) {\n+\t\t\tmy $written = syswrite($out, $data,\n+\t\t\t\tlength($data) - $offset, $offset);\n+\t\t\tdefined($written) && $written or die \"write response: $!\";\n+\t\t\t$offset += $written;\n+\t\t}\n+\t}\n+\n+\tsub start_response {\n+\t\tmy $out = $server->accept() or die \"accept: $!\";\n+\t\t<$out> or die \"read request: $!\";\n+\t\tmy $start = 0;\n+\t\twhile (<$out>) {\n+\t\t\tlast if /^\\r?\\n$/;\n+\t\t\t$start = $1 if /^Range: bytes=(\\d+)-/i;\n+\t\t}\n+\t\t$start < length($pack) or die \"invalid range $start\";\n+\t\tmy $length = length($pack) - $start;\n+\t\tmy $middle = int($length / 2);\n+\t\tmy $status = $start ? \"206 Partial Content\" : \"200 OK\";\n+\t\tmy $headers = \"HTTP/1.1 $status\\r\\n\" .\n+\t\t\t\"Content-Length: $length\\r\\n\" .\n+\t\t\t($start ? \"Content-Range: bytes $start-\" .\n+\t\t\t\t(length($pack) - 1) . \"/\" . length($pack) . \"\\r\\n\" : \"\") .\n+\t\t\t\"Connection: close\\r\\n\\r\\n\";\n+\t\twrite_all($out, $headers);\n+\t\twrite_all($out, substr($pack, $start, $middle));\n+\t\treturn ($out, $start + $middle);\n+\t}\n+\n+\tsignal_ready($server_ready, $server->sockport());\n+\tmy ($first, $first_pos) = start_response();\n+\tsignal_ready($first_ready, \"ready\");\n+\tmy ($second, $second_pos) = start_response();\n+\twrite_all($first, substr($pack, $first_pos));\n+\twrite_all($second, substr($pack, $second_pos));\n+\tclose($first) or die \"close first response: $!\";\n+\tclose($second) or die \"close second response: $!\";\n+\t$completed = 1;\n+\talarm 0;\n+\tEOF\n+\t{\n+\t\t\"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n+\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n+\t\t\t\"$TRASH_DIRECTORY/first-ready\" >server.log 2>&1 &\n+\t\tserver_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $server_pid 2>/dev/null || :\n+\t\twait $server_pid 2>/dev/null || :\n+\t\texec 7>&-\n+\t\texec 8>&-\n+\t\trm -f server-ready first-ready slow-pack-server\n+\t\" &&\n+\tread port <&7 &&\n+\turl=\"http://127.0.0.1:$port/pack\" &&\n+\t{\n+\t\t(\n+\t\t\tif ! GIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-first.trace\" \\\n+\t\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$url\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\ttest_path_is_file \"$tmpfile\" &&\n+\t{\n+\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n+\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\"$url\" >second.out &\n+\t\tsecond_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $second_pid 2>/dev/null || :\n+\t\twait $second_pid 2>/dev/null || :\n+\t\" &&\n+\twait \"$second_pid\" &&\n+\twait \"$first_pid\" &&\n+\twait \"$server_pid\" &&\n+\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n+\tsort first.out second.out >actual &&\n+\ttest_cmp expect actual &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-overlap cat-file -e \"$blob\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548983","messageId":"fee6f292cba09a8190cc78595d0aa80e31243a8d.1785047139.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785047139.git.tnyman@openai.com","subject":"[PATCH v5 3/3] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-26T06:44:48Z","receivedAt":"2026-07-26T06:44:54Z","isPatch":true,"body":"When index-pack finds an existing keep file it reports pack rather than\nkeep. Accept either result from http-fetch, and only register a keep\nlockfile when this fetch created it.\n\nRead the pack/keep prefix and hash without consuming any following fsck\noutput, validate the reported pack hash against the advertised hash, and\nexercise a packfile URI fetch with a pre-existing keep file.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n fetch-pack.c           | 33 ++++++++++++++++++---------------\n t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 15 deletions(-)\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 29c41132ee..e9f24fbd63 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t}\n \n \tfor (i = 0; i < packfile_uris.nr; i++) {\n+\t\tbool created_keep;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 1];\n+\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n-\t\t    memcmp(packname, \"keep\\t\", 5))\n-\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n+\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n+\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n+\t\t     memcmp(packhash, \"pack\\t\", 5)))\n+\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n+\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n \n-\t\tif (read_in_full(cmd.out, packname,\n-\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n-\t\t    packname[the_hash_algo->hexsz] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n-\n-\t\tpackname[the_hash_algo->hexsz] = '\\0';\n+\t\tif (read_in_full(cmd.out, packhash,\n+\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n+\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n+\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n+\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n \n \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n \n@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (finish_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to finish http-fetch\");\n \n-\t\tif (memcmp(packfile_uris.items[i].string, packname,\n+\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n \t\t\t   the_hash_algo->hexsz))\n \t\t\tdie(\"fetch-pack: pack downloaded from %s does not match expected hash %.*s\",\n \t\t\t    uri, (int) the_hash_algo->hexsz,\n \t\t\t    packfile_uris.items[i].string);\n \n-\t\tstring_list_append_nodup(pack_lockfiles,\n-\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n-\t\t\t\t\t\t repo_get_object_directory(the_repository),\n-\t\t\t\t\t\t packname));\n+\t\tif (created_keep)\n+\t\t\tstring_list_append_nodup(pack_lockfiles,\n+\t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n+\t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n+\t\t\t\t\t\t\t packhash));\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\ndiff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh\nindex 74a2b7730b..0f05286de8 100755\n--- a/t/t5702-protocol-v2.sh\n+++ b/t/t5702-protocol-v2.sh\n@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '\n \t\tfetch \"$HTTPD_URL/smart/http_parent\"\n '\n \n+test_expect_success 'packfile URI preserves an existing keep file' '\n+\tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n+\trm -rf \"$P\" http_child keep.expect &&\n+\n+\tgit init \"$P\" &&\n+\tgit -C \"$P\" config uploadpack.allowsidebandall true &&\n+\n+\techo my-blob >\"$P/my-blob\" &&\n+\tgit -C \"$P\" add my-blob &&\n+\tgit -C \"$P\" commit -m x &&\n+\tconfigure_exclusion \"$P\" my-blob >h &&\n+\n+\tgit init http_child &&\n+\tpackhash=$(cat packh) &&\n+\tkeep=\"http_child/.git/objects/pack/pack-$packhash.keep\" &&\n+\techo pre-existing >\"$keep\" &&\n+\tcp \"$keep\" keep.expect &&\n+\n+\tGIT_TEST_SIDEBAND_ALL=1 \\\n+\tgit -C http_child -c protocol.version=2 \\\n+\t\t-c fetch.uriprotocols=http,https \\\n+\t\tfetch \"$HTTPD_URL/smart/http_parent\" &&\n+\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.pack\" &&\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.idx\" &&\n+\ttest_cmp keep.expect \"$keep\" &&\n+\tgit -C http_child cat-file -e \"$(cat h)\"\n+'\n+\n test_expect_success 'fetching with valid packfile URI but invalid hash fails' '\n \tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n \trm -rf \"$P\" http_child log &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"548997","messageId":"20260726092027.GA3529827@coredump.intra.peff.net","threadId":"65988","inReplyTo":"d9667c93b03d1a71df55a33f90538b31afd08677.1785047139.git.tnyman@openai.com","subject":"Re: [PATCH v5 2/3] http: avoid concurrent appends to partial packs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-26T09:20:27Z","receivedAt":"2026-07-26T09:20:28Z","isPatch":true,"body":"On Sat, Jul 25, 2026 at 11:44:47PM -0700, Ted Nyman wrote:\n\n> Pack requests stage downloads in a predictable partial-pack file so an\n> interrupted transfer can be resumed. Both packfile URI and ordinary dumb\n> HTTP requests use this staging path. Opening it in append mode forces\n> each write to the current end of the file, so concurrent responses can\n> append duplicate data and corrupt the pack.\n> \n> Open the partial pack read-write without O_APPEND and seek once to its\n> current end. Each downloader then retains the offset matching the Range\n> it requested. Because the staging key must uniquely identify immutable\n> pack contents, overlapping responses write the same bytes at the same\n> offsets instead of extending the file with duplicate data.\n\nOK. I still think this is kind of horrible and gross, but I can't think\nof a reason it won't work (at least on POSIX-ish systems) and it solves\nthe problem with minimal changes and risk of regression.\n\nI wondered about racing with another concurrent writer on the seek, but\nI think it is OK. We seek immediately, and then use that offset (which\nwe get from another seek, replacing ftell()) as the value for our range\nrequest. So even if somebody else advances the file, we have _some_\natomic value that we'll start writing to ourselves, and the worst case\nis redundantly requesting a few bytes.\n\n> MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n> existing file. Create a missing partial pack exclusively, close it, and\n> reopen it without O_CREAT so every retained descriptor permits another\n> downloader to unlink the staging path. Duplicate that descriptor for\n> index-pack instead of reopening the path after closing the stream;\n> index-pack installs its own pack and the shared staging file is only\n> unlinked, never renamed.\n\nThis part I have no real knowledge or opinion on the Windows bits (or if\nthere's an easier way to do it).\n\n> Accept HTTP 416 when a partial pack is already\n> complete and let index-pack validate its contents.\n\nI wonder if we still need this or not. AIUI the original 416 responses\ncame because we were asking for nonsense outside of the range (because\nthe corrupted writes advanced the file too far). The worst case now is\nthat we'd ask for bytes \"N-\" when the file is only N bytes long, and the\nserver should say \"OK, here are your 0 bytes\". But maybe there's a\nserver who complains about that.\n\n> diff --git a/http.c b/http.c\n> index caccf2108e..a0d399b274 100644\n> --- a/http.c\n> +++ b/http.c\n> @@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)\n>  \tint tmpfile_fd;\n>  \tint ret = 0;\n>  \n> +\t/* Another downloader may unlink the staging path while we index it. */\n> +\ttmpfile_fd = xdup(fileno(preq->packfile));\n>  \tfclose(preq->packfile);\n>  \tpreq->packfile = NULL;\n> -\n> -\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n> +\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n> +\t\tdie_errno(\"unable to seek local file %s for pack\",\n> +\t\t\t  preq->tmpfile.buf);\n\nOK, here we are avoiding the race that it gets unlinked by dup-ing the\nexisting descriptor and seeking back to the start. Makes sense. But\nthen...\n\n> @@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n>  \telse\n>  \t\tip.no_stdout = 1;\n>  \n> -\tif (run_command(&ip)) {\n> +\tif (run_command(&ip))\n>  \t\tret = -1;\n> -\t\tgoto cleanup;\n> -\t}\n> -\n> -cleanup:\n> -\tclose(tmpfile_fd);\n>  \tunlink(preq->tmpfile.buf);\n>  \treturn ret;\n\nWhat is going on with this hunk? We don't really need to jump to cleanup\nhere because we get there directly anyway, and there are no other users\nof the cleanup label. So that part doesn't seem wrong, but rather\nunrelated.\n\nMore importantly, why don't we need to close tmpfile_fd anymore? We hand\nit off to run_command(), which will always close it. So I _think_ it was\nalways wrong to close it ourselves here. If so, then could this hunk\nbecome a preparatory commit on its own?\n\nThis commit is already confusing enough that the more extraneous stuff\nwe can take out of it the better.\n\n> @@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(\n> [...]\n> +\t/*\n> +\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n> +\t * existing file; reopen a newly created file so others may unlink it.\n> +\t */\n> +\tfor (;;) {\n> +\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n> +\t\tif (fd >= 0 || errno != ENOENT)\n> +\t\t\tbreak;\n> +\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n> +\t\tif (fd >= 0) {\n> +\t\t\tclose(fd);\n> +\t\t\tcontinue;\n> +\t\t}\n> +\t\tif (errno != EEXIST)\n> +\t\t\tbreak;\n> +\t}\n\nOK, and this is the opening magic. What's going on with the O_EXCL here,\nthough? We try to open once, and if that fails with ENOENT then we open\nagain. But isn't that racy? Two processes simultaneously try to open(),\nfind the file is not there, and then both try O_EXCL. Only one of them\nwill win, and the other will barf.\n\nI guess that is the reason for the loop, where we will try again over\nand over until we either pick up somebody else's copy or get our own.\nAnd if we get our own, we still close it and try again. And that's the\nWindows magic described in the commit message.\n\nThat is...subtle as hell. I really wonder if it would be worth\nintroducing the basic form of this (just opening once with O_RDWR) and\nthen doing the Windows hackery on top as a separate commit. That would\nleave the intermediate state subject to racy problems on Windows. But\nwhen balancing bisectability versus having a clear human-readable patch,\nI think I'd rather see it broken up.\n\n> +\tif (fd < 0) {\n> +\t\terror_errno(\"unable to open local file %s for pack\",\n> +\t\t\t    preq->tmpfile.buf);\n>  \t\tgoto abort;\n>  \t}\n\nOK, and then we get here if we broke out of the loop due to an error\nbesides ENOENT/EEXIST.\n\n> +\tprev_posn = lseek(fd, 0, SEEK_END);\n> +\tif (prev_posn < 0) {\n> +\t\terror_errno(\"unable to seek local file %s for pack\",\n> +\t\t\t    preq->tmpfile.buf);\n> +\t\tclose(fd);\n> +\t\tgoto abort;\n> +\t}\n> +\tpreq->packfile = xfdopen(fd, \"w\");\n\nAnd then this is the positioning magic to replace O_APPEND. Good.\n\n> diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\n\nFor the record, I don't love that we are using a custom perl script here\ninstead of going through apache (like all of our other tests). But I\nsuspect the apache version would be sufficiently horrific (possibly even\nworse) that it's not really worth pursuing. Hopefully this perl script\n(and the accompanying fifo monstrosities) can sit here for eternity\nun-looked-at by human eyes, just quietly doing their job until the heat\ndeath of the universe.\n\n-Peff\n"},{"id":"548998","messageId":"20260726092153.GB3529827@coredump.intra.peff.net","threadId":"65988","inReplyTo":"cover.1785047139.git.tnyman@openai.com","subject":"Re: [PATCH v5 0/3] packfile URIs: support concurrent downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-26T09:21:53Z","receivedAt":"2026-07-26T09:21:55Z","isPatch":true,"body":"On Sat, Jul 25, 2026 at 11:44:45PM -0700, Ted Nyman wrote:\n\n> Changes since v4:\n> \n>   * Clarify that the first --index-pack-arg specifies the command and\n>     subsequent instances specify its arguments.\n>   * Drop assumptions about which concurrent response reaches the\n>     staging file first. Either write order exercises the same\n>     overlapping-download behavior.\n>   * No production code changes.\n\nThanks. I hadn't really reviewed the code in v4 carefully, but I did so\nfor v5. I _think_ it is all correct, but it there are a few confusing\nbits in the middle patch that might be worth breaking apart for\nreadability.\n\nI could live with it as-is, though.\n\n-Peff\n"},{"id":"548999","messageId":"20260726100421.12648-1-tnyman@openai.com","threadId":"65988","inReplyTo":"20260726092027.GA3529827@coredump.intra.peff.net","subject":"Re: [PATCH v5 2/3] http: avoid concurrent appends to partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-26T10:04:21Z","receivedAt":"2026-07-26T10:04:24Z","isPatch":true,"body":"On Sun, Jul 26, 2026 at 05:20:27AM -0400, Jeff King wrote:\n> I wonder if we still need this or not.\n\nI think so, but wouldn't bet the farm on it. A concurrent downloader can\ncomplete the staging file before another downloader issues its Range\nrequest. That request then starts exactly at EOF, so the server can\nrespond with 416. The existing regression test exercises that case, and\nwe still need to let index-pack validate the completed local pack.\n\n> More importantly, why don't we need to close tmpfile_fd anymore? We hand\n> it off to run_command(), which will always close it. So I _think_ it was\n> always wrong to close it ourselves here. If so, then could this hunk\n> become a preparatory commit on its own?\n\nYou're right: run_command() already closes ip.in, so the old\nclose(tmpfile_fd) was a double-close. That cleanup is independent, and\nI can pull it into a preparatory patch if that would make the series\neasier to follow.\n\n> That is...subtle as hell. I really wonder if it would be worth\n> introducing the basic form of this (just opening once with O_RDWR) and\n> then doing the Windows hackery on top as a separate commit.\n\nI'm certainly not an expert on the Windows side, so I had to track this\ndown in the MinGW open() wrapper. The existing-file O_RDWR path includes\nFILE_SHARE_DELETE, while creating a new file falls back to _wopen()\nwithout it. The loop creates the file with O_EXCL if needed, closes that\ndescriptor, and retries through the existing-file path; a racing creator\nthat sees EEXIST also retries.\n\nI kept those pieces together to avoid an intermediate state without the\nrequired sharing behavior on MinGW, but I'm happy to split them if you\nthink it would be clearer.\n\n> Hopefully this perl script (and the accompanying fifo monstrosities)\n> can sit here for eternity un-looked-at by human eyes, just quietly\n> doing their job until the heat death of the universe.\n\nI thought you, of all people, might appreciate a little more Perl. ;-)\n\nThanks,\nTed\n"},{"id":"549000","messageId":"20260726102712.GA3535709@coredump.intra.peff.net","threadId":"65988","inReplyTo":"20260726100421.12648-1-tnyman@openai.com","subject":"Re: [PATCH v5 2/3] http: avoid concurrent appends to partial packs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-26T10:27:12Z","receivedAt":"2026-07-26T10:27:14Z","isPatch":true,"body":"On Sun, Jul 26, 2026 at 03:04:21AM -0700, Ted Nyman wrote:\n\n> On Sun, Jul 26, 2026 at 05:20:27AM -0400, Jeff King wrote:\n> > I wonder if we still need this or not.\n> \n> I think so, but wouldn't bet the farm on it. A concurrent downloader can\n> complete the staging file before another downloader issues its Range\n> request. That request then starts exactly at EOF, so the server can\n> respond with 416. The existing regression test exercises that case, and\n> we still need to let index-pack validate the completed local pack.\n\nYeah, the big question there for me is whether a server would return a\n416 in such a case. It seems reasonable that a client might know it has\nN bytes but not the full size, and ask for \"N-\", expecting to get some\nequivalent of a 0-byte read(). Whereas a 416 does not make it clear at\nall whether the range is nonsense, or if you happened to be at EOF.\n\nBut sadly we do not seem to live in that world, based on a few tests. So\nI agree we do need it to cover that edge case.\n\nIf you are splitting things out of the patch, can we do the same for\nthis 416 handling? It is already a problem even without concurrency if\nyou happen to get the full file but then fail for other reasons before\nindexing the pack (transient system errors, etc).\n\nTo be clear, I can live with things as they are and I don't want to make\ntoo much work for you in splitting. But I'm hoping that feeding it to an\nelectronic friend could do that split without much effort.\n\n> > That is...subtle as hell. I really wonder if it would be worth\n> > introducing the basic form of this (just opening once with O_RDWR) and\n> > then doing the Windows hackery on top as a separate commit.\n> \n> I'm certainly not an expert on the Windows side, so I had to track this\n> down in the MinGW open() wrapper. The existing-file O_RDWR path includes\n> FILE_SHARE_DELETE, while creating a new file falls back to _wopen()\n> without it. The loop creates the file with O_EXCL if needed, closes that\n> descriptor, and retries through the existing-file path; a racing creator\n> that sees EEXIST also retries.\n\nYeah, I understand it now after reading the commit message and the code\nseveral times. The loop is what I think is subtle, but I can't see a\nmore obvious way of writing it that deals with all of the possible\ncombinations and races.\n\n> I kept those pieces together to avoid an intermediate state without the\n> required sharing behavior on MinGW, but I'm happy to split them if you\n> think it would be clearer.\n\nIMHO it is the lesser of two evils. Not because we won't eventually end\nup with the subtle loop, but because it helps make the desired change in\nthe \"simpler\" version much easier to see. IOW, the Windows patch is\nalways going to be confusing, but we can at least salvage the original.\n\n> > Hopefully this perl script (and the accompanying fifo monstrosities)\n> > can sit here for eternity un-looked-at by human eyes, just quietly\n> > doing their job until the heat death of the universe.\n> \n> I thought you, of all people, might appreciate a little more Perl. ;-)\n\nWell, if we have to write something like that, obviously Perl is the\nright choice. ;) It's more the custom socket handling. E.g., there are\nsometimes subtle issues around listen-port allocation, especially when\nwe run the same test concurrently with --stress. Your script solves it\nby asking for a dynamic port and then passing that back to the caller\nover the fifo. That should work reliably, I think, it's just not our\nusual solution (but usually we are constrained to having to tell apache\nthe correct port up front, which means we need to pick an unambiguous\none ourselves).\n\n-Peff\n"},{"id":"549053","messageId":"cover.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785047139.git.tnyman@openai.com","subject":"[PATCH v6 0/6] packfile URIs: support concurrent downloads","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:37Z","receivedAt":"2026-07-27T00:28:56Z","isPatch":true,"body":"Packfile URI and dumb HTTP downloads stage packs at\nobjects/pack/pack-<hash>.pack.temp so an interrupted transfer can\nresume. Opening that file in append mode forces every write to its\ncurrent end. Two Git processes fetching the same pack into one object\ndatabase can therefore append duplicate data and corrupt the pack.\n\nThe first patch separates the unrelated --index-pack-arg documentation\nand error-message correction requested during review.\n\nThe second patch fixes an existing double-close when\nfinish_http_pack_request() passes its staging-file descriptor to\nindex-pack. start_command() already takes ownership of that descriptor,\nincluding when starting the child fails.\n\nThe third patch handles a completed partial pack independently of\nconcurrent downloads. A previous attempt can finish the transfer but\nfail before indexing it; retrying then requests a range starting at EOF.\nServers may respond with HTTP 416 in that case. Treat the response as a\ncompleted download and let index-pack validate the pack.\n\nThe fourth patch keeps the predictable staging name but removes append\nmode. Each downloader seeks once to the current end, requests the\ncorresponding Range, and writes using its own descriptor offset. Since\nthe staging key must identify immutable pack contents, overlapping\nresponses write identical bytes at identical offsets. There is no need\nfor pwrite(2) or cross-process coordination, and resumption continues to\nwork for both packfile URI and ordinary dumb HTTP downloads.\n\nThe fifth patch handles the additional MinGW sharing requirement. Its\nnon-append O_RDWR open grants FILE_SHARE_DELETE only for an existing\nfile. Create a missing staging file exclusively, close it, and reopen\nit without O_CREAT so every retained descriptor permits another\ndownloader to unlink the path.\n\nThe final patch handles the related .keep race. When another process has\nalready created the keep file, index-pack reports \"pack<TAB><hash>\"\ninstead of \"keep<TAB><hash>\". Accept both successful forms and remove\nonly keep files created by the current process. Read only the prefix and\nhash so any following fsck output remains available to fetch-pack.\n\nThe tests cover resumption, a completed partial returning 416,\noverlapping downloads, unlinking the staging path while index-pack holds\nits descriptor, and a pre-existing .keep file. The completed-partial and\nunlink tests do not require FIFOs, so they can run on MinGW even though\nthe concurrent-download test is skipped there.\n\nChanges since v5:\n\n* Split the existing double-close fix, HTTP 416 handling, generic\n  concurrent-download fix, and Windows sharing fix into separate\n  patches.\n* Replace the FIFO-based concurrent HTTP 416 test with a standalone\n  completed-partial test. Besides simplifying the test, this covers the\n  non-concurrent interrupted-download case directly.\n* Keep the final production code unchanged.\n\nEach patch passes t5550-http-fetch-dumb.sh. The final series also passes\nt5702-protocol-v2.sh, and the overlapping-download test passes 240 runs\nwith 12 parallel stress jobs.\n\nThe v5 discussion is at:\n\nhttps://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/\n\nTed Nyman (6):\n  http-fetch: correct --index-pack-arg documentation\n  http: avoid closing index-pack input twice\n  http: accept HTTP 416 for complete partial packs\n  http: avoid concurrent appends to partial packs\n  http: permit unlinking partial packs on Windows\n  fetch-pack: accept \"pack\" output for packfile URIs\n\n Documentation/git-http-fetch.adoc |  14 +-\n fetch-pack.c                      |  33 ++---\n http-fetch.c                      |   7 +-\n http-push.c                       |   3 +-\n http-walker.c                     |   3 +-\n http.c                            |  56 +++++---\n t/t5550-http-fetch-dumb.sh        | 204 ++++++++++++++++++++++++++++++\n t/t5702-protocol-v2.sh            |  31 +++++\n 8 files changed, 305 insertions(+), 46 deletions(-)\n\nRange-diff against v5:\n1:  a79af009ea = 1:  b5050a88ca http-fetch: correct --index-pack-arg documentation\n-:  ---------- > 2:  28662b0fd8 http: avoid closing index-pack input twice\n-:  ---------- > 3:  677e5399eb http: accept HTTP 416 for complete partial packs\n2:  d9667c93b0 ! 4:  7a83eb7091 http: avoid concurrent appends to partial packs\n    @@ Commit message\n         pack contents, overlapping responses write the same bytes at the same\n         offsets instead of extending the file with duplicate data.\n     \n    -    MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n    -    existing file. Create a missing partial pack exclusively, close it, and\n    -    reopen it without O_CREAT so every retained descriptor permits another\n    -    downloader to unlink the staging path. Duplicate that descriptor for\n    -    index-pack instead of reopening the path after closing the stream;\n    -    index-pack installs its own pack and the shared staging file is only\n    -    unlinked, never renamed. Accept HTTP 416 when a partial pack is already\n    -    complete and let index-pack validate its contents.\n    +    Duplicate the staging descriptor for index-pack instead of reopening the\n    +    path after closing the stream. Another downloader may unlink the staging\n    +    path before indexing begins, but index-pack can still read the retained\n    +    descriptor.\n     \n    -    Exercise resumed transfers, EOF ranges, overlapping 200 and 206\n    -    responses, and unlinking the staging path while index-pack still holds\n    -    its descriptor. Clarify the staging-key documentation.\n    +    Exercise resumed transfers and overlapping 200 and 206 responses, and\n    +    clarify the staging-key documentation.\n     \n         Signed-off-by: Ted Nyman <tnyman@openai.com>\n     \n    @@ Documentation/git-http-fetch.adoc: commit-id::\n      \n      --index-pack-arg=<arg>::\n     \n    - ## http-fetch.c ##\n    -@@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,\n    - \n    - \tif (start_active_slot(preq->slot)) {\n    - \t\trun_active_slot(preq->slot);\n    --\t\tif (results.curl_result != CURLE_OK) {\n    -+\t\tif (results.curl_result != CURLE_OK &&\n    -+\t\t    results.http_code != 416) {\n    - \t\t\tstruct url_info url;\n    - \t\t\tchar *nurl = url_normalize(preq->url, &url);\n    - \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\n    -\n    - ## http-push.c ##\n    -@@ http-push.c: static void finish_request(struct transfer_request *request)\n    - \n    - \t} else if (request->state == RUN_FETCH_PACKED) {\n    - \t\tint fail = 1;\n    --\t\tif (request->curl_result != CURLE_OK) {\n    -+\t\tif (request->curl_result != CURLE_OK &&\n    -+\t\t    request->http_code != 416) {\n    - \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n    - \t\t\t\trequest->url, curl_errorstr);\n    - \t\t} else {\n    -\n    - ## http-walker.c ##\n    -@@ http-walker.c: static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n    - \n    - \tif (start_active_slot(preq->slot)) {\n    - \t\trun_active_slot(preq->slot);\n    --\t\tif (results.curl_result != CURLE_OK) {\n    -+\t\tif (results.curl_result != CURLE_OK &&\n    -+\t\t    results.http_code != 416) {\n    - \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n    - \t\t\t      curl_errorstr);\n    - \t\t\tgoto abort;\n    -\n      ## http.c ##\n     @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)\n      \tint tmpfile_fd;\n    @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)\n      \n      \tip.git_cmd = 1;\n      \tip.in = tmpfile_fd;\n    -@@ http.c: int finish_http_pack_request(struct http_pack_request *preq)\n    - \telse\n    - \t\tip.no_stdout = 1;\n    - \n    --\tif (run_command(&ip)) {\n    -+\tif (run_command(&ip))\n    - \t\tret = -1;\n    --\t\tgoto cleanup;\n    --\t}\n    --\n    --cleanup:\n    --\tclose(tmpfile_fd);\n    - \tunlink(preq->tmpfile.buf);\n    - \treturn ret;\n    - }\n     @@ http.c: struct http_pack_request *new_http_pack_request(\n      struct http_pack_request *new_direct_http_pack_request(\n      \tconst unsigned char *packed_git_hash, char *url)\n    @@ http.c: struct http_pack_request *new_http_pack_request(\n     -\tif (!preq->packfile) {\n     -\t\terror(\"Unable to open local file %s for pack\",\n     -\t\t      preq->tmpfile.buf);\n    -+\t/*\n    -+\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n    -+\t * existing file; reopen a newly created file so others may unlink it.\n    -+\t */\n    -+\tfor (;;) {\n    -+\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n    -+\t\tif (fd >= 0 || errno != ENOENT)\n    -+\t\t\tbreak;\n    -+\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n    -+\t\tif (fd >= 0) {\n    -+\t\t\tclose(fd);\n    -+\t\t\tcontinue;\n    -+\t\t}\n    -+\t\tif (errno != EEXIST)\n    -+\t\t\tbreak;\n    -+\t}\n    ++\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT, 0666);\n     +\tif (fd < 0) {\n     +\t\terror_errno(\"unable to open local file %s for pack\",\n     +\t\t\t    preq->tmpfile.buf);\n    - \t\tgoto abort;\n    - \t}\n    ++\t\tgoto abort;\n    ++\t}\n     +\tprev_posn = lseek(fd, 0, SEEK_END);\n     +\tif (prev_posn < 0) {\n     +\t\terror_errno(\"unable to seek local file %s for pack\",\n     +\t\t\t    preq->tmpfile.buf);\n     +\t\tclose(fd);\n    -+\t\tgoto abort;\n    -+\t}\n    + \t\tgoto abort;\n    + \t}\n     +\tpreq->packfile = xfdopen(fd, \"w\");\n      \n      \tpreq->slot = get_active_slot();\n    @@ http.c: struct http_pack_request *new_direct_http_pack_request(\n      \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\n     \n      ## t/t5550-http-fetch-dumb.sh ##\n    -@@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n    - \tgit -C packfileclient cat-file -e \"$HASH\"\n    +@@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile accepts an already complete partial'\n    + \tgit -C packfileclient-complete cat-file -e \"$HASH\"\n      '\n      \n     +test_expect_success 'http-fetch --packfile resumes a partial download' '\n    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '\n     +\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n     +'\n     +\n    -+test_expect_success 'http-fetch --packfile permits unlink while indexing' '\n    -+\tgit init packfileclient-unlink &&\n    -+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n    -+\t\tls objects/pack/pack-*.pack) &&\n    -+\ttmpfile=\"packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n    -+\twrite_script git-unlink-index-pack <<-\\EOF &&\n    -+\ttest -f \"$GIT_TEST_PACK_TEMP\" || exit 1\n    -+\trm \"$GIT_TEST_PACK_TEMP\" || exit 1\n    -+\texec git index-pack \"$@\"\n    -+\tEOF\n    -+\ttest_when_finished \"rm -f git-unlink-index-pack\" &&\n    -+\tPATH=\"$TRASH_DIRECTORY:$PATH\" \\\n    -+\tGIT_TEST_PACK_TEMP=\"$TRASH_DIRECTORY/$tmpfile\" \\\n    -+\tgit -C packfileclient-unlink http-fetch --packfile=\"$ARBITRARY\" \\\n    -+\t\t--index-pack-arg=unlink-index-pack \\\n    -+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n    -+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n    -+\ttest_path_is_missing \"$tmpfile\" &&\n    -+\tgit -C packfileclient-unlink cat-file -e \"$HASH\"\n    -+'\n    -+\n    -+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '\n    -+\tgit init packfileclient-concurrent &&\n    -+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n    -+\t\tls objects/pack/pack-*.pack) &&\n    -+\tpackhash=$(basename \"$p\" .pack) &&\n    -+\tpackhash=${packhash#pack-} &&\n    -+\ttmpfile=\"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp\" &&\n    -+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n    -+\tmkfifo first-ready first-continue &&\n    -+\texec 8<>first-ready &&\n    -+\texec 9<>first-continue &&\n    -+\twrite_script git-wait-index-pack <<-\\EOF &&\n    -+\techo ready >\"$GIT_TEST_WAIT_READY\" &&\n    -+\tread continue <\"$GIT_TEST_WAIT_CONTINUE\" &&\n    -+\texec git index-pack \"$@\"\n    -+\tEOF\n    -+\t{\n    -+\t\t(\n    -+\t\t\tif ! PATH=\"$TRASH_DIRECTORY:$PATH\" \\\n    -+\t\t\tGIT_TEST_WAIT_READY=\"$TRASH_DIRECTORY/first-ready\" \\\n    -+\t\t\tGIT_TEST_WAIT_CONTINUE=\"$TRASH_DIRECTORY/first-continue\" \\\n    -+\t\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/first.trace\" \\\n    -+\t\t\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n    -+\t\t\t\t--index-pack-arg=wait-index-pack \\\n    -+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n    -+\t\t\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >first.out\n    -+\t\t\tthen\n    -+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n    -+\t\t\t\texit 1\n    -+\t\t\tfi\n    -+\t\t) &\n    -+\t\tfirst_pid=$!\n    -+\t} &&\n    -+\ttest_when_finished \"\n    -+\t\techo continue >&9\n    -+\t\tkill $first_pid 2>/dev/null || :\n    -+\t\twait $first_pid 2>/dev/null || :\n    -+\t\texec 8>&-\n    -+\t\texec 9>&-\n    -+\t\trm -f first-ready first-continue git-wait-index-pack\n    -+\t\" &&\n    -+\tread ready <&8 &&\n    -+\ttest \"$ready\" = ready &&\n    -+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/second.trace\" \\\n    -+\tgit -C packfileclient-concurrent http-fetch --packfile=\"$packhash\" \\\n    -+\t\t--index-pack-arg=index-pack \\\n    -+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n    -+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >second.out &&\n    -+\techo continue >&9 &&\n    -+\twait \"$first_pid\" &&\n    -+\tprintf \"pack\\t%s\\n\" \"$packhash\" >expect &&\n    -+\ttest_cmp expect first.out &&\n    -+\tprintf \"keep\\t%s\\n\" \"$packhash\" >expect &&\n    -+\ttest_cmp expect second.out &&\n    -+\ttest_grep \"Range: bytes=64-\" first.trace &&\n    -+\ttest_grep \"Range: bytes=[0-9]*-\" second.trace &&\n    -+\ttest_grep \"416 Requested Range Not Satisfiable\" second.trace &&\n    -+\ttest_path_is_missing \"$tmpfile\" &&\n    -+\tgit -C packfileclient-concurrent cat-file -e \"$HASH\"\n    -+'\n    -+\n     +test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n     +\tgit init packfileclient-overlap &&\n     +\tblob=$(test-tool genrandom pack-overlap 2m |\n-:  ---------- > 5:  87a20ac80f http: permit unlinking partial packs on Windows\n3:  fee6f292cb = 6:  be9e2fe273 fetch-pack: accept \"pack\" output for packfile URIs\n\nbase-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200\n-- \n2.55.0.openai.131.g83a728de1eb6\n"},{"id":"549054","messageId":"b5050a88ca8aa0b73c099ac773b07bd2bf398eb1.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"[PATCH v6 1/6] http-fetch: correct --index-pack-arg documentation","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:38Z","receivedAt":"2026-07-27T00:28:59Z","isPatch":true,"body":"The --packfile mode accepts one --index-pack-arg=<arg> option per\nargument passed to index-pack, but its documentation and option\ndependency errors still refer to the plural --index-pack-args form.\n\nCorrect the spelling and describe the repeatable per-argument form.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc | 9 +++++----\n http-fetch.c                      | 4 ++--\n 2 files changed, 7 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 2200f073c4..12036e65e9 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -50,11 +50,12 @@ commit-id::\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n \tThe hash is used to determine the name of the temporary file and is\n \tarbitrary. The output of index-pack is printed to stdout. Requires\n-\t--index-pack-args.\n+\tone or more --index-pack-arg options.\n \n---index-pack-args=<args>::\n-\tFor internal use only. The command to run on the contents of the\n-\tdownloaded pack. Arguments are URL-encoded separated by spaces.\n+--index-pack-arg=<arg>::\n+\tFor internal use only. The first instance specifies the command run on\n+\tthe contents of the downloaded pack. Subsequent instances specify its\n+\targuments.\n \n --recover::\n \tVerify that everything reachable from target is fetched.  Used after\ndiff --git a/http-fetch.c b/http-fetch.c\nindex f9b6ecb061..601a77c3c1 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)\n \n \tif (packfile) {\n \t\tif (!index_pack_args.nr)\n-\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-args\");\n+\t\t\tdie(_(\"the option '%s' requires '%s'\"), \"--packfile\", \"--index-pack-arg\");\n \n \t\tfetch_single_packfile(&packfile_hash, argv[arg],\n \t\t\t\t      index_pack_args.v);\n@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)\n \t}\n \n \tif (index_pack_args.nr)\n-\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-args\", \"--packfile\");\n+\t\tdie(_(\"the option '%s' requires '%s'\"), \"--index-pack-arg\", \"--packfile\");\n \n \tif (commits_on_stdin) {\n \t\tcommits = walker_targets_stdin(&commit_id, &write_ref);\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"549055","messageId":"28662b0fd892ecf6246be185ccb2d4654fb780a5.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"[PATCH v6 2/6] http: avoid closing index-pack input twice","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:39Z","receivedAt":"2026-07-27T00:29:01Z","isPatch":true,"body":"finish_http_pack_request() passes its staging-file descriptor to\nindex-pack through child_process.in. start_command() takes ownership\nof a supplied descriptor and closes it, even when starting the child\nfails.\n\nDo not close the descriptor again after run_command() returns.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n http.c | 7 +------\n 1 file changed, 1 insertion(+), 6 deletions(-)\n\ndiff --git a/http.c b/http.c\nindex caccf2108e..89a1ccc6d2 100644\n--- a/http.c\n+++ b/http.c\n@@ -2704,13 +2704,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \telse\n \t\tip.no_stdout = 1;\n \n-\tif (run_command(&ip)) {\n+\tif (run_command(&ip))\n \t\tret = -1;\n-\t\tgoto cleanup;\n-\t}\n-\n-cleanup:\n-\tclose(tmpfile_fd);\n \tunlink(preq->tmpfile.buf);\n \treturn ret;\n }\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"549056","messageId":"677e5399eb8ce260f6aa98d91b5b2634ff95e46c.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"[PATCH v6 3/6] http: accept HTTP 416 for complete partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:40Z","receivedAt":"2026-07-27T00:29:04Z","isPatch":true,"body":"A resumed pack request may already have all bytes of the remote pack.\nA server can respond to the resulting Range request with HTTP 416\ninstead of returning an empty response.\n\nAccept that response in each pack-download caller and let index-pack\nvalidate the completed staging file. This can happen without concurrent\ndownloads when a previous attempt completed the transfer but failed\nbefore indexing it.\n\nAdd a regression test that seeds a complete partial pack and checks that\nhttp-fetch indexes it after the server returns HTTP 416.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n http-fetch.c               |  3 ++-\n http-push.c                |  3 ++-\n http-walker.c              |  3 ++-\n t/t5550-http-fetch-dumb.sh | 19 +++++++++++++++++++\n 4 files changed, 25 insertions(+), 3 deletions(-)\n\ndiff --git a/http-fetch.c b/http-fetch.c\nindex 601a77c3c1..05f68f306a 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\tstruct url_info url;\n \t\t\tchar *nurl = url_normalize(preq->url, &url);\n \t\t\tif (!nurl || !git_env_bool(\"GIT_TRACE_REDACT\", 1)) {\ndiff --git a/http-push.c b/http-push.c\nindex 60f6f8f054..ef8abe3908 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)\n \n \t} else if (request->state == RUN_FETCH_PACKED) {\n \t\tint fail = 1;\n-\t\tif (request->curl_result != CURLE_OK) {\n+\t\tif (request->curl_result != CURLE_OK &&\n+\t\t    request->http_code != 416) {\n \t\t\tfprintf(stderr, \"Unable to get pack file %s\\n%s\",\n \t\t\t\trequest->url, curl_errorstr);\n \t\t} else {\ndiff --git a/http-walker.c b/http-walker.c\nindex b58a3b2a92..abafca84d6 100644\n--- a/http-walker.c\n+++ b/http-walker.c\n@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,\n \n \tif (start_active_slot(preq->slot)) {\n \t\trun_active_slot(preq->slot);\n-\t\tif (results.curl_result != CURLE_OK) {\n+\t\tif (results.curl_result != CURLE_OK &&\n+\t\t    results.http_code != 416) {\n \t\t\terror(\"Unable to get pack file %s\\n%s\", preq->url,\n \t\t\t      curl_errorstr);\n \t\t\tgoto abort;\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex f00eeae48f..698bbb3160 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -293,6 +293,25 @@ test_expect_success 'http-fetch --packfile' '\n \tgit -C packfileclient cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile accepts an already complete partial' '\n+\tgit init packfileclient-complete &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\tpackhash=$(basename \"$p\" .pack) &&\n+\tpackhash=${packhash#pack-} &&\n+\ttmpfile=\"packfileclient-complete/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tcp \"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" \"$tmpfile\" &&\n+\tchmod u+w \"$tmpfile\" &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/complete.trace\" \\\n+\tgit -C packfileclient-complete http-fetch --packfile=\"$packhash\" \\\n+\t\t--index-pack-arg=index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_grep \"416 Requested Range Not Satisfiable\" complete.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-complete cat-file -e \"$HASH\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"549057","messageId":"7a83eb7091473d12839f357212a224fd53b09cda.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"[PATCH v6 4/6] http: avoid concurrent appends to partial packs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:41Z","receivedAt":"2026-07-27T00:29:08Z","isPatch":true,"body":"Pack requests stage downloads in a predictable partial-pack file so an\ninterrupted transfer can be resumed. Both packfile URI and ordinary dumb\nHTTP requests use this staging path. Opening it in append mode forces\neach write to the current end of the file, so concurrent responses can\nappend duplicate data and corrupt the pack.\n\nOpen the partial pack read-write without O_APPEND and seek once to its\ncurrent end. Each downloader then retains the offset matching the Range\nit requested. Because the staging key must uniquely identify immutable\npack contents, overlapping responses write the same bytes at the same\noffsets instead of extending the file with duplicate data.\n\nDuplicate the staging descriptor for index-pack instead of reopening the\npath after closing the stream. Another downloader may unlink the staging\npath before indexing begins, but index-pack can still read the retained\ndescriptor.\n\nExercise resumed transfers and overlapping 200 and 206 responses, and\nclarify the staging-key documentation.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n Documentation/git-http-fetch.adoc |   5 +-\n http.c                            |  34 ++++---\n t/t5550-http-fetch-dumb.sh        | 164 ++++++++++++++++++++++++++++++\n 3 files changed, 187 insertions(+), 16 deletions(-)\n\ndiff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc\nindex 12036e65e9..45e0d3d07c 100644\n--- a/Documentation/git-http-fetch.adoc\n+++ b/Documentation/git-http-fetch.adoc\n@@ -48,8 +48,9 @@ commit-id::\n \tline (which is not expected in\n \tthis case), 'git http-fetch' fetches the packfile directly at the given\n \tURL and uses index-pack to generate corresponding .idx and .keep files.\n-\tThe hash is used to determine the name of the temporary file and is\n-\tarbitrary. The output of index-pack is printed to stdout. Requires\n+\tThe hash is used to determine the name of the temporary file. It need\n+\tnot be the pack hash, but it must uniquely identify the pack contents\n+\tfor resumption. The output of index-pack is printed to stdout. Requires\n \tone or more --index-pack-arg options.\n \n --index-pack-arg=<arg>::\ndiff --git a/http.c b/http.c\nindex 89a1ccc6d2..ad07ef3549 100644\n--- a/http.c\n+++ b/http.c\n@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)\n \tint tmpfile_fd;\n \tint ret = 0;\n \n+\t/* Another downloader may unlink the staging path while we index it. */\n+\ttmpfile_fd = xdup(fileno(preq->packfile));\n \tfclose(preq->packfile);\n \tpreq->packfile = NULL;\n-\n-\ttmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);\n+\tif (lseek(tmpfile_fd, 0, SEEK_SET) < 0)\n+\t\tdie_errno(\"unable to seek local file %s for pack\",\n+\t\t\t  preq->tmpfile.buf);\n \n \tip.git_cmd = 1;\n \tip.in = tmpfile_fd;\n@@ -2733,22 +2736,30 @@ struct http_pack_request *new_http_pack_request(\n struct http_pack_request *new_direct_http_pack_request(\n \tconst unsigned char *packed_git_hash, char *url)\n {\n-\toff_t prev_posn = 0;\n+\toff_t prev_posn;\n \tstruct http_pack_request *preq;\n+\tint fd;\n \n \tCALLOC_ARRAY(preq, 1);\n \tstrbuf_init(&preq->tmpfile, 0);\n-\n \tpreq->url = url;\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tpreq->packfile = fopen(preq->tmpfile.buf, \"a\");\n-\tif (!preq->packfile) {\n-\t\terror(\"Unable to open local file %s for pack\",\n-\t\t      preq->tmpfile.buf);\n+\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT, 0666);\n+\tif (fd < 0) {\n+\t\terror_errno(\"unable to open local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tgoto abort;\n+\t}\n+\tprev_posn = lseek(fd, 0, SEEK_END);\n+\tif (prev_posn < 0) {\n+\t\terror_errno(\"unable to seek local file %s for pack\",\n+\t\t\t    preq->tmpfile.buf);\n+\t\tclose(fd);\n \t\tgoto abort;\n \t}\n+\tpreq->packfile = xfdopen(fd, \"w\");\n \n \tpreq->slot = get_active_slot();\n \tpreq->headers = object_request_headers();\n@@ -2757,12 +2768,7 @@ struct http_pack_request *new_direct_http_pack_request(\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);\n \tcurl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);\n \n-\t/*\n-\t * If there is data present from a previous transfer attempt,\n-\t * resume where it left off\n-\t */\n-\tprev_posn = ftello(preq->packfile);\n-\tif (prev_posn>0) {\n+\tif (prev_posn > 0) {\n \t\tif (http_is_verbose)\n \t\t\tfprintf(stderr,\n \t\t\t\t\"Resuming fetch of pack %s at byte %\"PRIuMAX\"\\n\",\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex 698bbb3160..86b9d87ef5 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -312,6 +312,170 @@ test_expect_success 'http-fetch --packfile accepts an already complete partial'\n \tgit -C packfileclient-complete cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile resumes a partial download' '\n+\tgit init packfileclient-resume &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\ttest_copy_bytes 64 <\"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p\" >\"$tmpfile\" &&\n+\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/resume.trace\" \\\n+\tgit -C packfileclient-resume http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=index-pack --index-pack-arg=--stdin \\\n+\t\t--index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_grep \"Range: bytes=64-\" resume.trace &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-resume cat-file -e \"$HASH\"\n+'\n+\n+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n+\tgit init packfileclient-overlap &&\n+\tblob=$(test-tool genrandom pack-overlap 2m |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\thash-object -w --stdin) &&\n+\tpackhash=$(printf \"%s\\n\" \"$blob\" |\n+\t\tgit -C \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \\\n+\t\t\tpack-objects \"$TRASH_DIRECTORY/overlap-pack\") &&\n+\tpack=\"$TRASH_DIRECTORY/overlap-pack-$packhash.pack\" &&\n+\ttmpfile=\"packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp\" &&\n+\tmkfifo server-ready first-ready &&\n+\texec 7<>server-ready &&\n+\texec 8<>first-ready &&\n+\twrite_script slow-pack-server \"$PERL_PATH\" <<-\\EOF &&\n+\tuse strict;\n+\tuse warnings;\n+\tuse IO::Socket::INET;\n+\n+\tmy ($packfile, $server_ready, $first_ready) = @ARGV;\n+\tmy $completed = 0;\n+\tEND {\n+\t\tif (!$completed) {\n+\t\t\tsignal_ready($server_ready, \"failed\");\n+\t\t\tsignal_ready($first_ready, \"failed\");\n+\t\t}\n+\t}\n+\n+\t$SIG{ALRM} = sub { die \"timed out serving concurrent pack requests\\n\" };\n+\talarm 60;\n+\n+\topen(my $in, \"<:raw\", $packfile) or die \"open $packfile: $!\";\n+\tmy $pack = do { local $/; <$in> };\n+\tclose($in) or die \"close $packfile: $!\";\n+\tmy $server = IO::Socket::INET->new(LocalAddr => \"127.0.0.1\",\n+\t\tLocalPort => 0, Proto => \"tcp\", Listen => 2, ReuseAddr => 1)\n+\t\tor die \"listen: $!\";\n+\n+\tsub signal_ready {\n+\t\tmy ($file, $value) = @_;\n+\t\topen(my $out, \">\", $file) or die \"open $file: $!\";\n+\t\tprint $out \"$value\\n\" or die \"write $file: $!\";\n+\t\tclose($out) or die \"close $file: $!\";\n+\t}\n+\n+\tsub write_all {\n+\t\tmy ($out, $data) = @_;\n+\t\tmy $offset = 0;\n+\t\twhile ($offset < length($data)) {\n+\t\t\tmy $written = syswrite($out, $data,\n+\t\t\t\tlength($data) - $offset, $offset);\n+\t\t\tdefined($written) && $written or die \"write response: $!\";\n+\t\t\t$offset += $written;\n+\t\t}\n+\t}\n+\n+\tsub start_response {\n+\t\tmy $out = $server->accept() or die \"accept: $!\";\n+\t\t<$out> or die \"read request: $!\";\n+\t\tmy $start = 0;\n+\t\twhile (<$out>) {\n+\t\t\tlast if /^\\r?\\n$/;\n+\t\t\t$start = $1 if /^Range: bytes=(\\d+)-/i;\n+\t\t}\n+\t\t$start < length($pack) or die \"invalid range $start\";\n+\t\tmy $length = length($pack) - $start;\n+\t\tmy $middle = int($length / 2);\n+\t\tmy $status = $start ? \"206 Partial Content\" : \"200 OK\";\n+\t\tmy $headers = \"HTTP/1.1 $status\\r\\n\" .\n+\t\t\t\"Content-Length: $length\\r\\n\" .\n+\t\t\t($start ? \"Content-Range: bytes $start-\" .\n+\t\t\t\t(length($pack) - 1) . \"/\" . length($pack) . \"\\r\\n\" : \"\") .\n+\t\t\t\"Connection: close\\r\\n\\r\\n\";\n+\t\twrite_all($out, $headers);\n+\t\twrite_all($out, substr($pack, $start, $middle));\n+\t\treturn ($out, $start + $middle);\n+\t}\n+\n+\tsignal_ready($server_ready, $server->sockport());\n+\tmy ($first, $first_pos) = start_response();\n+\tsignal_ready($first_ready, \"ready\");\n+\tmy ($second, $second_pos) = start_response();\n+\twrite_all($first, substr($pack, $first_pos));\n+\twrite_all($second, substr($pack, $second_pos));\n+\tclose($first) or die \"close first response: $!\";\n+\tclose($second) or die \"close second response: $!\";\n+\t$completed = 1;\n+\talarm 0;\n+\tEOF\n+\t{\n+\t\t\"$TRASH_DIRECTORY/slow-pack-server\" \"$pack\" \\\n+\t\t\t\"$TRASH_DIRECTORY/server-ready\" \\\n+\t\t\t\"$TRASH_DIRECTORY/first-ready\" >server.log 2>&1 &\n+\t\tserver_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $server_pid 2>/dev/null || :\n+\t\twait $server_pid 2>/dev/null || :\n+\t\texec 7>&-\n+\t\texec 8>&-\n+\t\trm -f server-ready first-ready slow-pack-server\n+\t\" &&\n+\tread port <&7 &&\n+\turl=\"http://127.0.0.1:$port/pack\" &&\n+\t{\n+\t\t(\n+\t\t\tif ! GIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-first.trace\" \\\n+\t\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\t\"$url\" >first.out\n+\t\t\tthen\n+\t\t\t\techo failed >\"$TRASH_DIRECTORY/first-ready\" &&\n+\t\t\t\texit 1\n+\t\t\tfi\n+\t\t) &\n+\t\tfirst_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $first_pid 2>/dev/null || :\n+\t\twait $first_pid 2>/dev/null || :\n+\t\" &&\n+\tread ready <&8 &&\n+\ttest \"$ready\" = ready &&\n+\ttest_path_is_file \"$tmpfile\" &&\n+\t{\n+\t\tGIT_TRACE_CURL=\"$TRASH_DIRECTORY/overlap-second.trace\" \\\n+\t\tGIT_TRACE_CURL_NO_DATA=1 \\\n+\t\tgit -C packfileclient-overlap http-fetch --packfile=\"$packhash\" \\\n+\t\t\t--index-pack-arg=index-pack \\\n+\t\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\t\"$url\" >second.out &\n+\t\tsecond_pid=$!\n+\t} &&\n+\ttest_when_finished \"\n+\t\tkill $second_pid 2>/dev/null || :\n+\t\twait $second_pid 2>/dev/null || :\n+\t\" &&\n+\twait \"$second_pid\" &&\n+\twait \"$first_pid\" &&\n+\twait \"$server_pid\" &&\n+\tprintf \"keep\\t%s\\npack\\t%s\\n\" \"$packhash\" \"$packhash\" | sort >expect &&\n+\tsort first.out second.out >actual &&\n+\ttest_cmp expect actual &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-overlap cat-file -e \"$blob\"\n+'\n+\n test_expect_success 'fetch notices corrupt pack' '\n \tcp -R \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n \t(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_bad1.git &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"549058","messageId":"87a20ac80fcaaaa3f69a24a650f5c135b24956ff.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"[PATCH v6 5/6] http: permit unlinking partial packs on Windows","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:42Z","receivedAt":"2026-07-27T00:29:10Z","isPatch":true,"body":"On Windows, an open file must permit FILE_SHARE_DELETE before another\nprocess can unlink it. MinGW's non-append O_RDWR open enables that\nsharing mode only for an existing file; adding O_CREAT falls back to\n_wopen(), which cannot set it.\n\nFirst try opening the partial pack without O_CREAT. If it does not\nexist, create it exclusively, close that descriptor, and retry through\nthe existing-file path. A racing creator retries after EEXIST.\n\nThis ensures that every retained descriptor permits another downloader\nto unlink the staging path. Add an unlink-while-indexing test that does\nnot require FIFOs and can therefore run on MinGW.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n http.c                     | 17 ++++++++++++++++-\n t/t5550-http-fetch-dumb.sh | 21 +++++++++++++++++++++\n 2 files changed, 37 insertions(+), 1 deletion(-)\n\ndiff --git a/http.c b/http.c\nindex ad07ef3549..a0d399b274 100644\n--- a/http.c\n+++ b/http.c\n@@ -2746,7 +2746,22 @@ struct http_pack_request *new_direct_http_pack_request(\n \n \todb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, \"pack\");\n \tstrbuf_addstr(&preq->tmpfile, \".temp\");\n-\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT, 0666);\n+\t/*\n+\t * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an\n+\t * existing file; reopen a newly created file so others may unlink it.\n+\t */\n+\tfor (;;) {\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR);\n+\t\tif (fd >= 0 || errno != ENOENT)\n+\t\t\tbreak;\n+\t\tfd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);\n+\t\tif (fd >= 0) {\n+\t\t\tclose(fd);\n+\t\t\tcontinue;\n+\t\t}\n+\t\tif (errno != EEXIST)\n+\t\t\tbreak;\n+\t}\n \tif (fd < 0) {\n \t\terror_errno(\"unable to open local file %s for pack\",\n \t\t\t    preq->tmpfile.buf);\ndiff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh\nindex 86b9d87ef5..b5758f1c9c 100755\n--- a/t/t5550-http-fetch-dumb.sh\n+++ b/t/t5550-http-fetch-dumb.sh\n@@ -328,6 +328,27 @@ test_expect_success 'http-fetch --packfile resumes a partial download' '\n \tgit -C packfileclient-resume cat-file -e \"$HASH\"\n '\n \n+test_expect_success 'http-fetch --packfile permits unlink while indexing' '\n+\tgit init packfileclient-unlink &&\n+\tp=$(cd \"$HTTPD_DOCUMENT_ROOT_PATH\"/repo_pack.git &&\n+\t\tls objects/pack/pack-*.pack) &&\n+\ttmpfile=\"packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp\" &&\n+\twrite_script git-unlink-index-pack <<-\\EOF &&\n+\ttest -f \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\trm \"$GIT_TEST_PACK_TEMP\" || exit 1\n+\texec git index-pack \"$@\"\n+\tEOF\n+\ttest_when_finished \"rm -f git-unlink-index-pack\" &&\n+\tPATH=\"$TRASH_DIRECTORY:$PATH\" \\\n+\tGIT_TEST_PACK_TEMP=\"$TRASH_DIRECTORY/$tmpfile\" \\\n+\tgit -C packfileclient-unlink http-fetch --packfile=\"$ARBITRARY\" \\\n+\t\t--index-pack-arg=unlink-index-pack \\\n+\t\t--index-pack-arg=--stdin --index-pack-arg=--keep \\\n+\t\t\"$HTTPD_URL/dumb/repo_pack.git/$p\" >out &&\n+\ttest_path_is_missing \"$tmpfile\" &&\n+\tgit -C packfileclient-unlink cat-file -e \"$HASH\"\n+'\n+\n test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '\n \tgit init packfileclient-overlap &&\n \tblob=$(test-tool genrandom pack-overlap 2m |\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"549059","messageId":"be9e2fe2735124ee16e2c87fc3aaf3d37dde416b.1785111375.git.tnyman@openai.com","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"[PATCH v6 6/6] fetch-pack: accept \"pack\" output for packfile URIs","fromName":"Ted Nyman","fromEmail":"tnyman@openai.com","sentAt":"2026-07-27T00:28:43Z","receivedAt":"2026-07-27T00:29:13Z","isPatch":true,"body":"When index-pack finds an existing keep file it reports pack rather than\nkeep. Accept either result from http-fetch, and only register a keep\nlockfile when this fetch created it.\n\nRead the pack/keep prefix and hash without consuming any following fsck\noutput, validate the reported pack hash against the advertised hash, and\nexercise a packfile URI fetch with a pre-existing keep file.\n\nSigned-off-by: Ted Nyman <tnyman@openai.com>\n---\n fetch-pack.c           | 33 ++++++++++++++++++---------------\n t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++\n 2 files changed, 49 insertions(+), 15 deletions(-)\n\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 29c41132ee..e9f24fbd63 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t}\n \n \tfor (i = 0; i < packfile_uris.nr; i++) {\n+\t\tbool created_keep;\n \t\tint j;\n \t\tstruct child_process cmd = CHILD_PROCESS_INIT;\n-\t\tchar packname[GIT_MAX_HEXSZ + 1];\n+\t\tchar packhash[GIT_MAX_HEXSZ + 1];\n \t\tconst char *uri = packfile_uris.items[i].string +\n \t\t\tthe_hash_algo->hexsz + 1;\n \n@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (start_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to spawn http-fetch\");\n \n-\t\tif (read_in_full(cmd.out, packname, 5) < 0 ||\n-\t\t    memcmp(packname, \"keep\\t\", 5))\n-\t\t\tdie(\"fetch-pack: expected keep then TAB at start of http-fetch output\");\n+\t\tif (read_in_full(cmd.out, packhash, 5) != 5 ||\n+\t\t    (memcmp(packhash, \"keep\\t\", 5) &&\n+\t\t     memcmp(packhash, \"pack\\t\", 5)))\n+\t\t\tdie(\"fetch-pack: expected pack or keep then TAB at start of http-fetch output\");\n+\t\tcreated_keep = !memcmp(packhash, \"keep\\t\", 5);\n \n-\t\tif (read_in_full(cmd.out, packname,\n-\t\t\t\t the_hash_algo->hexsz + 1) < 0 ||\n-\t\t    packname[the_hash_algo->hexsz] != '\\n')\n-\t\t\tdie(\"fetch-pack: expected hash then LF at end of http-fetch output\");\n-\n-\t\tpackname[the_hash_algo->hexsz] = '\\0';\n+\t\tif (read_in_full(cmd.out, packhash,\n+\t\t\t\t the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||\n+\t\t    packhash[the_hash_algo->hexsz] != '\\n')\n+\t\t\tdie(\"fetch-pack: expected hash then LF in http-fetch output\");\n+\t\tpackhash[the_hash_algo->hexsz] = '\\0';\n \n \t\tparse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);\n \n@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,\n \t\tif (finish_command(&cmd))\n \t\t\tdie(\"fetch-pack: unable to finish http-fetch\");\n \n-\t\tif (memcmp(packfile_uris.items[i].string, packname,\n+\t\tif (memcmp(packfile_uris.items[i].string, packhash,\n \t\t\t   the_hash_algo->hexsz))\n \t\t\tdie(\"fetch-pack: pack downloaded from %s does not match expected hash %.*s\",\n \t\t\t    uri, (int) the_hash_algo->hexsz,\n \t\t\t    packfile_uris.items[i].string);\n \n-\t\tstring_list_append_nodup(pack_lockfiles,\n-\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n-\t\t\t\t\t\t repo_get_object_directory(the_repository),\n-\t\t\t\t\t\t packname));\n+\t\tif (created_keep)\n+\t\t\tstring_list_append_nodup(pack_lockfiles,\n+\t\t\t\t\t\t xstrfmt(\"%s/pack/pack-%s.keep\",\n+\t\t\t\t\t\t\t repo_get_object_directory(the_repository),\n+\t\t\t\t\t\t\t packhash));\n \t}\n \tstring_list_clear(&packfile_uris, 0);\n \tstrvec_clear(&index_pack_args);\ndiff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh\nindex 74a2b7730b..0f05286de8 100755\n--- a/t/t5702-protocol-v2.sh\n+++ b/t/t5702-protocol-v2.sh\n@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '\n \t\tfetch \"$HTTPD_URL/smart/http_parent\"\n '\n \n+test_expect_success 'packfile URI preserves an existing keep file' '\n+\tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n+\trm -rf \"$P\" http_child keep.expect &&\n+\n+\tgit init \"$P\" &&\n+\tgit -C \"$P\" config uploadpack.allowsidebandall true &&\n+\n+\techo my-blob >\"$P/my-blob\" &&\n+\tgit -C \"$P\" add my-blob &&\n+\tgit -C \"$P\" commit -m x &&\n+\tconfigure_exclusion \"$P\" my-blob >h &&\n+\n+\tgit init http_child &&\n+\tpackhash=$(cat packh) &&\n+\tkeep=\"http_child/.git/objects/pack/pack-$packhash.keep\" &&\n+\techo pre-existing >\"$keep\" &&\n+\tcp \"$keep\" keep.expect &&\n+\n+\tGIT_TEST_SIDEBAND_ALL=1 \\\n+\tgit -C http_child -c protocol.version=2 \\\n+\t\t-c fetch.uriprotocols=http,https \\\n+\t\tfetch \"$HTTPD_URL/smart/http_parent\" &&\n+\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.pack\" &&\n+\ttest_path_is_file \\\n+\t\t\"http_child/.git/objects/pack/pack-$packhash.idx\" &&\n+\ttest_cmp keep.expect \"$keep\" &&\n+\tgit -C http_child cat-file -e \"$(cat h)\"\n+'\n+\n test_expect_success 'fetching with valid packfile URI but invalid hash fails' '\n \tP=\"$HTTPD_DOCUMENT_ROOT_PATH/http_parent\" &&\n \trm -rf \"$P\" http_child log &&\n-- \n2.55.0.openai.131.g83a728de1eb6\n\n"},{"id":"549237","messageId":"xmqqcxw5o4m8.fsf@gitster.g","threadId":"65988","inReplyTo":"cover.1785111375.git.tnyman@openai.com","subject":"Re: [PATCH v6 0/6] packfile URIs: support concurrent downloads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-29T21:41:51Z","receivedAt":"2026-07-29T21:41:55Z","isPatch":true,"body":"Ted Nyman <tnyman@openai.com> writes:\n\n> Changes since v5:\n>\n> * Split the existing double-close fix, HTTP 416 handling, generic\n>   concurrent-download fix, and Windows sharing fix into separate\n>   patches.\n> * Replace the FIFO-based concurrent HTTP 416 test with a standalone\n>   completed-partial test. Besides simplifying the test, this covers the\n>   non-concurrent interrupted-download case directly.\n> * Keep the final production code unchanged.\n>\n> Each patch passes t5550-http-fetch-dumb.sh. The final series also passes\n> t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs\n> with 12 parallel stress jobs.\n>\n> The v5 discussion is at:\n>\n> https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/\n\nIs everybody happy with this new iteration?\n\nThe design of the re-download feature itself, as far as I\nunderstand, was favourably accepted from the earliest iteration, and\nnow the CI breakages were corrected with the latest iteration of the\ntests, so we should be in pretty good shape, I presume.\n\nThanks.\n\n\n"},{"id":"549390","messageId":"20260801135313.GA2041176@coredump.intra.peff.net","threadId":"65988","inReplyTo":"28662b0fd892ecf6246be185ccb2d4654fb780a5.1785111375.git.tnyman@openai.com","subject":"Re: [PATCH v6 2/6] http: avoid closing index-pack input twice","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-08-01T13:53:13Z","receivedAt":"2026-08-01T13:53:21Z","isPatch":true,"body":"On Sun, Jul 26, 2026 at 05:28:39PM -0700, Ted Nyman wrote:\n\n> finish_http_pack_request() passes its staging-file descriptor to\n> index-pack through child_process.in. start_command() takes ownership\n> of a supplied descriptor and closes it, even when starting the child\n> fails.\n> \n> Do not close the descriptor again after run_command() returns.\n\nThanks for splitting this out.\n\n> @@ -2704,13 +2704,8 @@ int finish_http_pack_request(struct http_pack_request *preq)\n>  \telse\n>  \t\tip.no_stdout = 1;\n>  \n> -\tif (run_command(&ip)) {\n> +\tif (run_command(&ip))\n>  \t\tret = -1;\n> -\t\tgoto cleanup;\n> -\t}\n> -\n> -cleanup:\n> -\tclose(tmpfile_fd);\n\nThe patch _could_ just be a one-liner dropping this close(). Removing\nthe cleanup label here is optional, but is a simplification that works\nbecause nobody else jumps to it (which must be true because we'd fail to\ncompile otherwise).\n\nI probably would have mentioned that in the commit message, but I think\nthere's diminishing returns in trying to polish further.\n\n-Peff\n"},{"id":"549391","messageId":"20260801135844.GB2041176@coredump.intra.peff.net","threadId":"65988","inReplyTo":"677e5399eb8ce260f6aa98d91b5b2634ff95e46c.1785111375.git.tnyman@openai.com","subject":"Re: [PATCH v6 3/6] http: accept HTTP 416 for complete partial packs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-08-01T13:58:44Z","receivedAt":"2026-08-01T13:58:45Z","isPatch":true,"body":"On Sun, Jul 26, 2026 at 05:28:40PM -0700, Ted Nyman wrote:\n\n> A resumed pack request may already have all bytes of the remote pack.\n> A server can respond to the resulting Range request with HTTP 416\n> instead of returning an empty response.\n> \n> Accept that response in each pack-download caller and let index-pack\n> validate the completed staging file. This can happen without concurrent\n> downloads when a previous attempt completed the transfer but failed\n> before indexing it.\n> \n> Add a regression test that seeds a complete partial pack and checks that\n> http-fetch indexes it after the server returns HTTP 416.\n\nAgain, thanks for splitting this out and demonstrating the\nnon-concurrent case. It all looks good to me.\n\nI do wonder what will happen when we get a 416 and we _don't_ have a\ncomplete pack. E.g., imagine the file size on the server changed (it\nshouldn't if they are using the hash of the pack contents as the name,\nbut that's not strictly required).\n\nPreviously we'd barf on the curl error. Now we'll guess that we got the\nfull file, even though we have a partial download. Presumably we'd then\njust barf at the index-pack level. I guess this is not really any\ndifferent than other resumption problems. If the file changed on the\nserver, we could easily download half of one version and half of\nanother. Ultimately we don't trust any of it until index-pack processes\nthe whole thing.\n\nSo this seems like a good direction to me.\n\n-Peff\n"},{"id":"549392","messageId":"20260801140255.GC2041176@coredump.intra.peff.net","threadId":"65988","inReplyTo":"xmqqcxw5o4m8.fsf@gitster.g","subject":"Re: [PATCH v6 0/6] packfile URIs: support concurrent downloads","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-08-01T14:02:55Z","receivedAt":"2026-08-01T14:02:57Z","isPatch":true,"body":"On Wed, Jul 29, 2026 at 02:41:51PM -0700, Junio C Hamano wrote:\n\n> Ted Nyman <tnyman@openai.com> writes:\n> \n> > Changes since v5:\n> >\n> > * Split the existing double-close fix, HTTP 416 handling, generic\n> >   concurrent-download fix, and Windows sharing fix into separate\n> >   patches.\n> > * Replace the FIFO-based concurrent HTTP 416 test with a standalone\n> >   completed-partial test. Besides simplifying the test, this covers the\n> >   non-concurrent interrupted-download case directly.\n> > * Keep the final production code unchanged.\n> >\n> > Each patch passes t5550-http-fetch-dumb.sh. The final series also passes\n> > t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs\n> > with 12 parallel stress jobs.\n> >\n> > The v5 discussion is at:\n> >\n> > https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/\n> \n> Is everybody happy with this new iteration?\n> \n> The design of the re-download feature itself, as far as I\n> understand, was favourably accepted from the earliest iteration, and\n> now the CI breakages were corrected with the latest iteration of the\n> tests, so we should be in pretty good shape, I presume.\n\nYeah, sorry, I hadn't had time to look carefully. I just did so, and it\nall looks good to me. v6 splits the patches in a way that (at least to\nmy mind) make the trickiest parts of the logic easier to follow.\n\n-Peff\n"},{"id":"550097","messageId":"xmqqik5kbmyh.fsf@gitster.g","threadId":"65988","inReplyTo":"20260801140255.GC2041176@coredump.intra.peff.net","subject":"Re: [PATCH v6 0/6] packfile URIs: support concurrent downloads","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-08T16:23:34Z","receivedAt":"2026-08-08T16:23:38Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> On Wed, Jul 29, 2026 at 02:41:51PM -0700, Junio C Hamano wrote:\n>\n>> Ted Nyman <tnyman@openai.com> writes:\n>> \n>> > Changes since v5:\n>> >\n>> > * Split the existing double-close fix, HTTP 416 handling, generic\n>> >   concurrent-download fix, and Windows sharing fix into separate\n>> >   patches.\n>> > * Replace the FIFO-based concurrent HTTP 416 test with a standalone\n>> >   completed-partial test. Besides simplifying the test, this covers the\n>> >   non-concurrent interrupted-download case directly.\n>> > * Keep the final production code unchanged.\n>> >\n>> > Each patch passes t5550-http-fetch-dumb.sh. The final series also passes\n>> > t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs\n>> > with 12 parallel stress jobs.\n>> >\n>> > The v5 discussion is at:\n>> >\n>> > https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/\n>> \n>> Is everybody happy with this new iteration?\n>> \n>> The design of the re-download feature itself, as far as I\n>> understand, was favourably accepted from the earliest iteration, and\n>> now the CI breakages were corrected with the latest iteration of the\n>> tests, so we should be in pretty good shape, I presume.\n>\n> Yeah, sorry, I hadn't had time to look carefully. I just did so, and it\n> all looks good to me. v6 splits the patches in a way that (at least to\n> my mind) make the trickiest parts of the logic easier to follow.\n\nThanks.\n\n"}]}