Volume XXII, number 279Tuesday, October 6, 2026Latest message 32 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

patch, 2 partshttp: use unique tempfiles for packfile URI downloads

57 messages between Jul 13, 2026 and Aug 8, 2026, from Ted Nyman, Junio C Hamano, Taylor Blau, Jeff King.

Plain Markdown or JSON for tools and agents. Diffs are folded; open one to read it.

Ted NymanJul 13, 2026, 22:34 UTC in reply to Ted Nyman on lore

Since 8d5d2a34df (http-fetch: support fetching packfiles by URL, 2020-06-10), packfile URI downloads have been staged at objects/pack/pack-<hash>.pack.temp.

The path is derived from the advertised pack hash. Two processes fetching the same pack into a shared object database therefore open the same file for append. Their writes can corrupt the temporary pack. If one process arrives after the other has completed the download, it may instead try to resume at EOF, which some HTTP servers reject with 416.

Use the tempfile API to give direct packfile URI downloads unique temporary files. Keep the deterministic path for ordinary dumb HTTP pack requests, which use it to resume a partial download left by an earlier invocation.

This means that a packfile URI download cannot be resumed by a later invocation. A retry starts with an empty temporary file instead.

Add a test which pauses one process after downloading the pack and starts another process using the same object database.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc |  5 +-
 http.c                            | 77 +++++++++++++++++++++----------
 http.h                            |  1 +
 t/t5550-http-fetch-dumb.sh        | 72 ++++++++++++++++++++++++++++-
 4 files changed, 126 insertions(+), 29 deletions(-)
Show changes to 4 files +126 −29

Documentation/git-http-fetch.adoc, http.c, http.h, t/t5550-http-fetch-dumb.sh

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 2200f073c4..533bf381c4 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -48,9 +48,8 @@ commit-id::
 	line (which is not expected in
 	this case), 'git http-fetch' fetches the packfile directly at the given
 	URL and uses index-pack to generate corresponding .idx and .keep files.
-	The hash is used to determine the name of the temporary file and is
-	arbitrary. The output of index-pack is printed to stdout. Requires
-	--index-pack-args.
+	The hash is arbitrary. The output of index-pack is printed to stdout.
+	Requires --index-pack-args.
 
 --index-pack-args=<args>::
 	For internal use only. The command to run on the contents of the
diff --git a/http.c b/http.c
index b4e7b8d00b..5a46e7c65c 100644
--- a/http.c
+++ b/http.c
@@ -2668,7 +2668,10 @@ int http_get_info_packs(const char *base_url, struct packfile_list *packs)
 
 void release_http_pack_request(struct http_pack_request *preq)
 {
-	if (preq->packfile) {
+	if (preq->tempfile) {
+		delete_tempfile(&preq->tempfile);
+		preq->packfile = NULL;
+	} else if (preq->packfile) {
 		fclose(preq->packfile);
 		preq->packfile = NULL;
 	}
@@ -2688,7 +2691,10 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	int tmpfile_fd;
 	int ret = 0;
 
-	fclose(preq->packfile);
+	if (preq->tempfile)
+		close_tempfile_gently(preq->tempfile);
+	else
+		fclose(preq->packfile);
 	preq->packfile = NULL;
 
 	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
@@ -2711,7 +2717,10 @@ int finish_http_pack_request(struct http_pack_request *preq)
 
 cleanup:
 	close(tmpfile_fd);
-	unlink(preq->tmpfile.buf);
+	if (preq->tempfile)
+		delete_tempfile(&preq->tempfile);
+	else
+		unlink(preq->tmpfile.buf);
 	return ret;
 }
 
@@ -2723,20 +2732,8 @@ void http_install_packfile(struct packed_git *p,
 	packfile_store_add_pack(files->packed, p);
 }
 
-struct http_pack_request *new_http_pack_request(
-	const unsigned char *packed_git_hash, const char *base_url) {
-
-	struct strbuf buf = STRBUF_INIT;
-
-	end_url_with_slash(&buf, base_url);
-	strbuf_addf(&buf, "objects/pack/pack-%s.pack",
-		hash_to_hex(packed_git_hash));
-	return new_direct_http_pack_request(packed_git_hash,
-					    strbuf_detach(&buf, NULL));
-}
-
-struct http_pack_request *new_direct_http_pack_request(
-	const unsigned char *packed_git_hash, char *url)
+static struct http_pack_request *new_http_pack_request_for_url(
+	const unsigned char *packed_git_hash, char *url, int resumable)
 {
 	off_t prev_posn = 0;
 	struct http_pack_request *preq;
@@ -2746,9 +2743,22 @@ struct http_pack_request *new_direct_http_pack_request(
 
 	preq->url = url;
 
-	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
-	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
+	if (resumable) {
+		odb_pack_name(the_repository, &preq->tmpfile,
+			      packed_git_hash, "pack");
+		strbuf_addstr(&preq->tmpfile, ".temp");
+		preq->packfile = fopen(preq->tmpfile.buf, "a");
+	} else {
+		strbuf_addf(&preq->tmpfile, "%s/pack/tmp_pack_XXXXXX",
+			    repo_get_object_directory(the_repository));
+		preq->tempfile = mks_tempfile_m(preq->tmpfile.buf, 0444);
+		if (preq->tempfile) {
+			strbuf_reset(&preq->tmpfile);
+			strbuf_addstr(&preq->tmpfile,
+				      get_tempfile_path(preq->tempfile));
+			preq->packfile = fdopen_tempfile(preq->tempfile, "w");
+		}
+	}
 	if (!preq->packfile) {
 		error("Unable to open local file %s for pack",
 		      preq->tmpfile.buf);
@@ -2766,8 +2776,9 @@ struct http_pack_request *new_direct_http_pack_request(
 	 * If there is data present from a previous transfer attempt,
 	 * resume where it left off
 	 */
-	prev_posn = ftello(preq->packfile);
-	if (prev_posn>0) {
+	if (resumable)
+		prev_posn = ftello(preq->packfile);
+	if (prev_posn > 0) {
 		if (http_is_verbose)
 			fprintf(stderr,
 				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
@@ -2779,12 +2790,28 @@ struct http_pack_request *new_direct_http_pack_request(
 	return preq;
 
 abort:
-	strbuf_release(&preq->tmpfile);
-	free(preq->url);
-	free(preq);
+	release_http_pack_request(preq);
 	return NULL;
 }
 
+struct http_pack_request *new_http_pack_request(
+	const unsigned char *packed_git_hash, const char *base_url)
+{
+	struct strbuf buf = STRBUF_INIT;
+
+	end_url_with_slash(&buf, base_url);
+	strbuf_addf(&buf, "objects/pack/pack-%s.pack",
+		hash_to_hex(packed_git_hash));
+	return new_http_pack_request_for_url(packed_git_hash,
+					     strbuf_detach(&buf, NULL), 1);
+}
+
+struct http_pack_request *new_direct_http_pack_request(
+	const unsigned char *packed_git_hash, char *url)
+{
+	return new_http_pack_request_for_url(packed_git_hash, url, 0);
+}
+
 /* Helpers for fetching objects (loose) */
 static size_t fwrite_sha1_file(char *ptr, size_t eltsize, size_t nmemb,
 			       void *data)
diff --git a/http.h b/http.h
index 729c51904d..2c900779f5 100644
--- a/http.h
+++ b/http.h
@@ -224,6 +224,7 @@ struct http_pack_request {
 
 	FILE *packfile;
 	struct strbuf tmpfile;
+	struct tempfile *tempfile;
 	struct active_request_slot *slot;
 	struct curl_slist *headers;
 };
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index b0080bf204..314a74c433 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -293,6 +293,74 @@ test_expect_success 'http-fetch --packfile' '
 	git -C packfileclient cat-file -e "$HASH"
 '
 
+test_expect_success PIPE 'concurrent http-fetch --packfile' '
+	git init packfileclient-concurrent &&
+	HASH=$(git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git rev-parse HEAD) &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	packhash=$(basename "$p" .pack) &&
+	packhash=${packhash#pack-} &&
+
+	mkfifo first-ready first-continue &&
+	exec 8<>first-ready &&
+	exec 9<>first-continue &&
+	write_script git-wait-index-pack <<-\EOF &&
+	echo ready >"$GIT_TEST_WAIT_READY" &&
+	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
+	exec git index-pack "$@"
+	EOF
+
+	# Hold the first download before it is indexed, so that the second
+	# download installs the pack first.
+	{
+		(
+			if ! PATH="$TRASH_DIRECTORY:$PATH" \
+			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
+			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
+			git -C packfileclient-concurrent http-fetch \
+				--packfile="$packhash" \
+				--index-pack-arg=wait-index-pack \
+				--index-pack-arg=--stdin \
+				--index-pack-arg=--keep \
+				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		echo continue >&9
+		wait $first_pid 2>/dev/null || :
+		exec 8>&-
+		exec 9>&-
+		rm -f first-ready first-continue git-wait-index-pack
+	" &&
+
+	read ready <&8 &&
+	test "$ready" = ready &&
+	git -C packfileclient-concurrent http-fetch \
+		--packfile="$packhash" \
+		--index-pack-arg=index-pack \
+		--index-pack-arg=--stdin \
+		--index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
+	echo continue >&9 &&
+	wait "$first_pid" &&
+
+	printf "pack\t%s\n" "$packhash" >expect &&
+	test_cmp expect first.out &&
+	printf "keep\t%s\n" "$packhash" >expect &&
+	test_cmp expect second.out &&
+	test_path_is_missing \
+		"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
+	find packfileclient-concurrent/.git/objects/pack \
+		-name "tmp_pack_*" -print >tmpfiles &&
+	test_must_be_empty tmpfiles &&
+	git -C packfileclient-concurrent cat-file -e "$HASH"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
@@ -313,7 +381,9 @@ test_expect_success 'http-fetch --packfile with corrupt pack' '
 	git init packfileclient &&
 	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git && ls objects/pack/pack-*.pack) &&
 	test_must_fail git -C packfileclient http-fetch --packfile \
-		"$HTTPD_URL"/dumb/repo_bad1.git/$p
+		"$HTTPD_URL"/dumb/repo_bad1.git/$p &&
+	find packfileclient/.git/objects/pack -name "tmp_pack_*" -print >tmpfiles &&
+	test_must_be_empty tmpfiles
 '
 
 test_expect_success 'fetch notices corrupt idx' '
-- 
2.55.0
Ted NymanJul 13, 2026, 22:34 UTC in reply to Ted Nyman on lore

[PATCH 2/2] fetch-pack: accept "pack" output for packfile URIs

When "index-pack --keep" creates a .keep file, it reports "keep<TAB><hash>". If the file already exists, index-pack leaves it untouched and reports "pack<TAB><hash>" instead.

Since dd4b732df7 (upload-pack: send part of packfile response as uri, 2020-06-10), fetch-pack has accepted only the "keep" form for packs downloaded through packfile URIs. A concurrent fetch can install the same pack and create its .keep file before another process reaches index-pack. The latter process then fails even though index-pack completed successfully.

Accept both successful forms. Add a path to pack_lockfiles only for the "keep" form, so cleanup removes only a keep file created by the current process and preserves a pre-existing one.

Add a regression test which pre-creates a keep file and verifies that a fetch succeeds without changing it.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 fetch-pack.c           | 36 ++++++++++++++++++++----------------
 t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
 2 files changed, 51 insertions(+), 16 deletions(-)
Show changes to 2 files +51 −16

fetch-pack.c, t/t5702-protocol-v2.sh

diff --git a/fetch-pack.c b/fetch-pack.c
index 120e01f3cf..a16b80177a 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1887,9 +1887,12 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 	}
 
 	for (i = 0; i < packfile_uris.nr; i++) {
+		int created_keep = 0;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 1];
+		char packname[GIT_MAX_HEXSZ + 6];
+		const char *packhash;
+		const int packname_len = the_hash_algo->hexsz + 6;
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1907,16 +1910,16 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, 5) < 0 ||
-		    memcmp(packname, "keep\t", 5))
-			die("fetch-pack: expected keep then TAB at start of http-fetch output");
-
-		if (read_in_full(cmd.out, packname,
-				 the_hash_algo->hexsz + 1) < 0 ||
-		    packname[the_hash_algo->hexsz] != '\n')
-			die("fetch-pack: expected hash then LF at end of http-fetch output");
-
-		packname[the_hash_algo->hexsz] = '\0';
+		if (read_in_full(cmd.out, packname, packname_len) != packname_len ||
+		    packname[packname_len - 1] != '\n')
+			die("fetch-pack: expected pack or keep, TAB, hash, "
+			    "then LF in http-fetch output");
+		packname[packname_len - 1] = '\0';
+		if (skip_prefix(packname, "keep\t", &packhash))
+			created_keep = 1;
+		else if (!skip_prefix(packname, "pack\t", &packhash))
+			die("fetch-pack: expected pack or keep, TAB, hash, "
+			    "then LF in http-fetch output");
 
 		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
 
@@ -1925,16 +1928,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (finish_command(&cmd))
 			die("fetch-pack: unable to finish http-fetch");
 
-		if (memcmp(packfile_uris.items[i].string, packname,
+		if (memcmp(packfile_uris.items[i].string, packhash,
 			   the_hash_algo->hexsz))
 			die("fetch-pack: pack downloaded from %s does not match expected hash %.*s",
 			    uri, (int) the_hash_algo->hexsz,
 			    packfile_uris.items[i].string);
 
-		string_list_append_nodup(pack_lockfiles,
-					 xstrfmt("%s/pack/pack-%s.keep",
-						 repo_get_object_directory(the_repository),
-						 packname));
+		if (created_keep)
+			string_list_append_nodup(pack_lockfiles,
+						 xstrfmt("%s/pack/pack-%s.keep",
+							 repo_get_object_directory(the_repository),
+							 packhash));
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);
diff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh
index 9f6cf4142d..1861eb7d7c 100755
--- a/t/t5702-protocol-v2.sh
+++ b/t/t5702-protocol-v2.sh
@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '
 		fetch "$HTTPD_URL/smart/http_parent"
 '
 
+test_expect_success 'packfile URI preserves an existing keep file' '
+	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
+	rm -rf "$P" http_child keep.expect &&
+
+	git init "$P" &&
+	git -C "$P" config uploadpack.allowsidebandall true &&
+
+	echo my-blob >"$P/my-blob" &&
+	git -C "$P" add my-blob &&
+	git -C "$P" commit -m x &&
+	configure_exclusion "$P" my-blob >h &&
+
+	git init http_child &&
+	packhash=$(cat packh) &&
+	keep="http_child/.git/objects/pack/pack-$packhash.keep" &&
+	echo pre-existing >"$keep" &&
+	cp "$keep" keep.expect &&
+
+	GIT_TEST_SIDEBAND_ALL=1 \
+	git -C http_child -c protocol.version=2 \
+		-c fetch.uriprotocols=http,https \
+		fetch "$HTTPD_URL/smart/http_parent" &&
+
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.pack" &&
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.idx" &&
+	test_cmp keep.expect "$keep" &&
+	git -C http_child cat-file -e "$(cat h)"
+'
+
 test_expect_success 'fetching with valid packfile URI but invalid hash fails' '
 	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
 	rm -rf "$P" http_child log &&
-- 
2.55.0
Ted NymanJul 13, 2026, 22:37 UTC on lore

[PATCH 0/2] packfile URIs: support concurrent downloads

Packfile URI downloads currently stage a pack at objects/pack/pack-<hash>.pack.temp. Two Git processes fetching the same pack into one object database can append to that file concurrently, which can corrupt the temporary pack or cause a resume request at EOF.

The first patch gives each direct packfile URI download a private temporary file. Ordinary dumb HTTP pack requests retain their existing resumable staging behavior. A later packfile URI retry starts a new download.

The second patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process.

Each patch adds a regression test for its respective race.
Ted Nyman (2):
  http: use unique tempfiles for packfile URI downloads
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  5 +-
 fetch-pack.c                      | 36 ++++++++-------
 http.c                            | 77 +++++++++++++++++++++----------
 http.h                            |  1 +
 t/t5550-http-fetch-dumb.sh        | 72 ++++++++++++++++++++++++++++-
 t/t5702-protocol-v2.sh            | 31 +++++++++++++
 6 files changed, 177 insertions(+), 45 deletions(-)
base-commit: e9019fcafe0040228b8631c30f97ae1adb61bcdc
-- 
2.55.0
Junio C HamanoJul 14, 2026, 01:00 UTC in reply to Ted Nyman on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

Ted Nyman <tnyman@openai.com> writes:
Show 17 quoted lines
> Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,
> 2020-06-10), packfile URI downloads have been staged at
> objects/pack/pack-<hash>.pack.temp.
>
> The path is derived from the advertised pack hash. Two processes
> fetching the same pack into a shared object database therefore open the
> same file for append. Their writes can corrupt the temporary pack. If
> one process arrives after the other has completed the download, it may
> instead try to resume at EOF, which some HTTP servers reject with 416.
>
> Use the tempfile API to give direct packfile URI downloads unique
> temporary files. Keep the deterministic path for ordinary dumb HTTP
> pack requests, which use it to resume a partial download left by an
> earlier invocation.
>
> This means that a packfile URI download cannot be resumed by a later
> invocation. A retry starts with an empty temporary file instead.

While that does sound like a safe and correct approach, stepping back briefly, would it not be wasteful for the second process to download the same packfile that the first has already started downloading?

Are there better ways for these processes to coordinate with each other? Instead of appending to the file, what if the second process uses a predictable temporary name (which we already use) to open a new file with O_CREAT | O_EXCL to avoid this redundant work? If the open call fails because the file already exists, the second process can detect that another process is active and wait for it to finish rather than initiating its own network request.

Doing so might require setting up a trigger or polling mechanism to wait for the first process's download to complete (and detecting if the other process dies without cleaning up), though that may open a can of worms.

Ted NymanJul 14, 2026, 01:58 UTC in reply to Junio C Hamano on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

> While that does sound like a safe and correct approach, stepping
> back briefly, would it not be wasteful for the second process to
> download the same packfile that the first has already started
> downloading?
Yes. If two fetches overlap, the second download is redundant.
> Are there better ways for these processes to coordinate with each
> other? Instead of appending to the file, what if the second process
> uses a predictable temporary name (which we already use) to open a
> new file with O_CREAT | O_EXCL to avoid this redundant work?

Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL would prevent concurrent writes, but EEXIST alone would not distinguish an in-progress download from one left by an earlier failed or interrupted invocation. The existing .pack.temp name is not covered by the tmp_* pruning path, so simply waiting for it to disappear could leave a fetch stuck after a crash.

The waiting case would also need a complete handoff. If the first process finishes, the second would need to notice the installed pack and account for the expected index-pack result and keep state. If the first process fails and removes its temporary file, the second would need to retry as the downloader. That is possible, but introduces cross-process coordination and a timeout policy in http-fetch.

The unique tempfile preserves the existing "download, index, then install" behavior for each invocation and fixes both the concurrent-append and EOF-resume failures. Avoiding the duplicate transfer would be useful for large packs, but I would prefer to keep that as a follow-up unless you think it is necessary for this correctness fix.

Thanks, Ted

Taylor BlauJul 14, 2026, 04:06 UTC in reply to Ted Nyman on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

On Mon, Jul 13, 2026 at 03:34:33PM -0700, Ted Nyman wrote:
Show 57 quoted lines
> Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,
> 2020-06-10), packfile URI downloads have been staged at
> objects/pack/pack-<hash>.pack.temp.
>
> The path is derived from the advertised pack hash. Two processes
> fetching the same pack into a shared object database therefore open the
> same file for append. Their writes can corrupt the temporary pack. If
> one process arrives after the other has completed the download, it may
> instead try to resume at EOF, which some HTTP servers reject with 416.
>
> Use the tempfile API to give direct packfile URI downloads unique
> temporary files. Keep the deterministic path for ordinary dumb HTTP
> pack requests, which use it to resume a partial download left by an
> earlier invocation.
>
> This means that a packfile URI download cannot be resumed by a later
> invocation. A retry starts with an empty temporary file instead.
>
> Add a test which pauses one process after downloading the pack and
> starts another process using the same object database.
>
> Signed-off-by: Ted Nyman <tnyman@openai.com>
> ---
>  Documentation/git-http-fetch.adoc |  5 +-
>  http.c                            | 77 +++++++++++++++++++++----------
>  http.h                            |  1 +
>  t/t5550-http-fetch-dumb.sh        | 72 ++++++++++++++++++++++++++++-
>  4 files changed, 126 insertions(+), 29 deletions(-)
>
> diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
> index 2200f073c4..533bf381c4 100644
> --- a/Documentation/git-http-fetch.adoc
> +++ b/Documentation/git-http-fetch.adoc
> @@ -48,9 +48,8 @@ commit-id::
>  	line (which is not expected in
>  	this case), 'git http-fetch' fetches the packfile directly at the given
>  	URL and uses index-pack to generate corresponding .idx and .keep files.
> -	The hash is used to determine the name of the temporary file and is
> -	arbitrary. The output of index-pack is printed to stdout. Requires
> -	--index-pack-args.
> +	The hash is arbitrary. The output of index-pack is printed to stdout.
> +	Requires --index-pack-args.
>
>  --index-pack-args=<args>::
>  	For internal use only. The command to run on the contents of the
> diff --git a/http.c b/http.c
> index b4e7b8d00b..5a46e7c65c 100644
> --- a/http.c
> +++ b/http.c
> @@ -2668,7 +2668,10 @@ int http_get_info_packs(const char *base_url, struct packfile_list *packs)
>
>  void release_http_pack_request(struct http_pack_request *preq)
>  {
> -	if (preq->packfile) {
> +	if (preq->tempfile) {
> +		delete_tempfile(&preq->tempfile);
> +		preq->packfile = NULL;

We should be able to drop the assignment to NULL on the second line, since `delete_tempfile()` takes a double pointer to the 'struct packfile' and NULL's it out for us.

(The other callers appear to avoid explicitly setting `preq->tempfile` to NULL.)

The rest of the patch looks good to me.
Show 21 quoted lines
> diff --git a/http.h b/http.h
> index 729c51904d..2c900779f5 100644
> --- a/http.h
> +++ b/http.h
> @@ -224,6 +224,7 @@ struct http_pack_request {
>
>  	FILE *packfile;
>  	struct strbuf tmpfile;
> +	struct tempfile *tempfile;
>  	struct active_request_slot *slot;
>  	struct curl_slist *headers;
>  };
> diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
> index b0080bf204..314a74c433 100755
> --- a/t/t5550-http-fetch-dumb.sh
> +++ b/t/t5550-http-fetch-dumb.sh
> @@ -293,6 +293,74 @@ test_expect_success 'http-fetch --packfile' '
>  	git -C packfileclient cat-file -e "$HASH"
>  '
>
> +test_expect_success PIPE 'concurrent http-fetch --packfile' '
Phew ;-).

This is definitely tricky to test, but what you wrote here looks plausibly correct to me.

Thanks, Taylor

Taylor BlauJul 14, 2026, 04:07 UTC in reply to Ted Nyman on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

On Mon, Jul 13, 2026 at 06:58:24PM -0700, Ted Nyman wrote:
Show 18 quoted lines
> > While that does sound like a safe and correct approach, stepping
> > back briefly, would it not be wasteful for the second process to
> > download the same packfile that the first has already started
> > downloading?
>
> Yes. If two fetches overlap, the second download is redundant.
>
> > Are there better ways for these processes to coordinate with each
> > other? Instead of appending to the file, what if the second process
> > uses a predictable temporary name (which we already use) to open a
> > new file with O_CREAT | O_EXCL to avoid this redundant work?
>
> Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL
> would prevent concurrent writes, but EEXIST alone would not
> distinguish an in-progress download from one left by an earlier
> failed or interrupted invocation. The existing .pack.temp name is not
> covered by the tmp_* pruning path, so simply waiting for it to
> disappear could leave a fetch stuck after a crash.

Exactly. If two processes are downloading the same pack at the same time to different locations, the effort is of course redundant. But I don't think we can reliably distinguish between that case and one where an earlier process died in the middle of downloading a pack but was unable to clean up after itself.

Thanks, Taylor

Taylor BlauJul 14, 2026, 04:13 UTC in reply to Ted Nyman on lore

Re: [PATCH 0/2] packfile URIs: support concurrent downloads

On Mon, Jul 13, 2026 at 03:37:58PM -0700, Ted Nyman wrote:
> Ted Nyman (2):
>   http: use unique tempfiles for packfile URI downloads
>   fetch-pack: accept "pack" output for packfile URIs

I left one pretty minor style-nit on the first patch, but otherwise this looks good to me.

Thanks, Taylor

Jeff KingJul 14, 2026, 05:28 UTC in reply to Ted Nyman on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

On Mon, Jul 13, 2026 at 06:58:24PM -0700, Ted Nyman wrote:
Show 11 quoted lines
> > Are there better ways for these processes to coordinate with each
> > other? Instead of appending to the file, what if the second process
> > uses a predictable temporary name (which we already use) to open a
> > new file with O_CREAT | O_EXCL to avoid this redundant work?
> 
> Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL
> would prevent concurrent writes, but EEXIST alone would not
> distinguish an in-progress download from one left by an earlier
> failed or interrupted invocation. The existing .pack.temp name is not
> covered by the tmp_* pruning path, so simply waiting for it to
> disappear could leave a fetch stuck after a crash.
A few thoughts:
  - Using O_EXCL makes this essentially a lockfile. So we could apply
    the logic used elsewhere for lockfiles, like auto-removing files
    with ancient mtimes. Or we could even go all-in with a pid check for
    liveness; most of Git's lockfiles don't do that, but at least one
    does (the background auto-gc lock).
  - If we're not already using a name which is auto-cleaned during
    maintenance, we probably ought to be. Leaving aside concurrency
    issues, nobody would ever clean up the on-disk cruft.
    But of course the original code here is intentionally _not_ using a
    name we'd clean up, because it wants to be able to resume an
    interrupted transfer.  And you're explicitly breaking that for the
    packfile URI case.
    Is that a cost we're OK with paying? Fixing it opens up that same
    coordination can of worms. You have to tell the difference a
    concurrent writer and a previous dead one (whose work you can
    resume).
    It does feel weird that we'd do one thing for dumb-http and another
    for packfile URIs. Wouldn't they suffer from the same concurrency
    and resumption problems?
Show 6 quoted lines
> The unique tempfile preserves the existing "download, index, then
> install" behavior for each invocation and fixes both the
> concurrent-append and EOF-resume failures. Avoiding the duplicate
> transfer would be useful for large packs, but I would prefer to keep
> that as a follow-up unless you think it is necessary for this
> correctness fix.

If we're OK with killing the ability to resume, then yeah, I think it would make sense to start simple and un-break things. And then put a coordination layer on top later (or never if nobody cares enough).

-Peff
Jeff KingJul 14, 2026, 05:44 UTC in reply to Taylor Blau on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

On Mon, Jul 13, 2026 at 09:06:12PM -0700, Taylor Blau wrote:
Show 13 quoted lines
> >  void release_http_pack_request(struct http_pack_request *preq)
> >  {
> > -	if (preq->packfile) {
> > +	if (preq->tempfile) {
> > +		delete_tempfile(&preq->tempfile);
> > +		preq->packfile = NULL;
> 
> We should be able to drop the assignment to NULL on the second line,
> since `delete_tempfile()` takes a double pointer to the 'struct
> packfile' and NULL's it out for us.
> 
> (The other callers appear to avoid explicitly setting `preq->tempfile`
> to NULL.)

It takes a double-pointer to the "struct tempfile"; the NULL assignment is to the "packfile" member, which is the FILE handle.

I thought at first this was buggy; we still call fdopen() on the tempfile and assign the result to preq->packfile, even in the new non-resumable case. Don't we need to fclose() it? But the answer is no: fdopen_tempfile() retains ownership of the result, storing it in tempfile.fp. So it will be correctly closed during delete_tempfile(), and in fact we must _not_ fclose it again.

But assigning NULL can happen with either style. So doing it unconditionally like:

  if (preq->tempfile)
	delete_tempfile(&preq->tempfile);
  else if (preq->packfile)
	fclose(preq->packfile);
  preq->packfile = NULL;

makes more sense, as it is done in finish_http_pack_request(). It might even make sense to add a comment explaining why we don't need to fclose() in the first part of the conditional.

All that said, I do not think setting it to NULL matters at all here, since the function ends with free(preq). So just dropping the NULL would perhaps be more clear.

-Peff
Jeff KingJul 14, 2026, 06:46 UTC in reply to Ted Nyman on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

On Mon, Jul 13, 2026 at 03:34:33PM -0700, Ted Nyman wrote:
Show 5 quoted lines
> The path is derived from the advertised pack hash. Two processes
> fetching the same pack into a shared object database therefore open the
> same file for append. Their writes can corrupt the temporary pack. If
> one process arrives after the other has completed the download, it may
> instead try to resume at EOF, which some HTTP servers reject with 416.

Yuck. In theory they're writing the same thing, but I think the source of the corruption is append mode. Two concurrent writers will keep auto-seeking to the end of the file, rather than keeping their own file pointers. There's no way to ask for O_APPEND without O_TRUNC via stdio, but we can drop down a level like this:

Show changes to http.c +4 −3
diff --git a/http.c b/http.c
index b4e7b8d00b..d7362c99a2 100644
--- a/http.c
+++ b/http.c
@@ -2740,6 +2740,7 @@ struct http_pack_request *new_direct_http_pack_request(
 {
 	off_t prev_posn = 0;
 	struct http_pack_request *preq;
+	int fd;
 
 	CALLOC_ARRAY(preq, 1);
 	strbuf_init(&preq->tmpfile, 0);
@@ -2748,12 +2749,13 @@ struct http_pack_request *new_direct_http_pack_request(
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
-	if (!preq->packfile) {
+	fd = open(preq->tmpfile.buf, O_WRONLY|O_CREAT, 0666);
+	if (fd < 0) {
 		error("Unable to open local file %s for pack",
 		      preq->tmpfile.buf);
 		goto abort;
 	}
+	preq->packfile = xfdopen(fd, "w");
 
 	preq->slot = get_active_slot();
 	preq->headers = object_request_headers();

That patch (with no other code changes) passes your test.

I suspect it could cause us to racily send an http range of "N-" to the
server, where N is the total number of bytes in the file (because we
don't know how many bytes there are supposed to be). I don't know if
that would cause an HTTP 416 or not. I think possibly not, and the 416
you saw (and that I see when running the test without any code changes)
might be from sending a range that starts _past_ N. We end up with a
too-long when both processes are appending.

I can't say I love the overall notion of "two processes are writing the
same data, it will probably be fine!". There might be portability
issues, and I'm not sure what would happen if we ever did get
conflicting data. If we're just feeding this to "index-pack --stdin"
we'd at least notice the problem (rather than quietly corrupting the
indexed file!).

So I'm offering this as a point for further discussion, and not
necessarily a counter-proposal. ;)

> Use the tempfile API to give direct packfile URI downloads unique
> temporary files. Keep the deterministic path for ordinary dumb HTTP
> pack requests, which use it to resume a partial download left by an
> earlier invocation.
> 
> This means that a packfile URI download cannot be resumed by a later
> invocation. A retry starts with an empty temporary file instead.

Arguably losing the ability to retry is a regression. In general, I
think we should prefer correctness to efficiency. But I wonder if this
is a case where the user might want to make the choice to say "I am not
going to fetch two packfiles at once; please enable resumable fetches".
Especially because one of the selling points of packfile URIs is that
they are resumable.

One other thought on resumable transfers: if we are not going to resume
the transfer, then why spool the pack to disk at all? In other words,
why not just send it straight to "index-pack --stdin". That fixes your
concurrency issue (because it uses its own tempfiles behind the scene),
but has two other big advantages:

  1. It halves the number of disk writes, and lowers the peak disk usage
     (with the current code, there is a moment where both the tempfile
     and the indexed pack are present on disk).

  2. It pipelines the data processing. The current code bottlenecks on
     the network while the CPU sits idle, and then bottlenecks on the
     CPU once we have the whole file. We could be doing useful CPU work
     during the network transfer, just like a regular pack code does.


So I'm not quite sold on losing the ability to resume entirely. And in
cases where we do lose it, I think it opens up other improvements.

But I'll reader over the rest of the patch with the notion that this is
the direction we want to go in.

> diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
> index 2200f073c4..533bf381c4 100644
> --- a/Documentation/git-http-fetch.adoc
> +++ b/Documentation/git-http-fetch.adoc
> @@ -48,9 +48,8 @@ commit-id::
>  	line (which is not expected in
>  	this case), 'git http-fetch' fetches the packfile directly at the given
>  	URL and uses index-pack to generate corresponding .idx and .keep files.
> -	The hash is used to determine the name of the temporary file and is
> -	arbitrary. The output of index-pack is printed to stdout. Requires
> -	--index-pack-args.
> +	The hash is arbitrary. The output of index-pack is printed to stdout.
> +	Requires --index-pack-args.

Do we even need to provide a hash anymore? After your patch I don't
think we even use it. It might be worth keeping around, though, as it
would be a unique key for de-duping or resuming, if we ever did
implement those on top.

>  void release_http_pack_request(struct http_pack_request *preq)
>  {
> -	if (preq->packfile) {
> +	if (preq->tempfile) {
> +		delete_tempfile(&preq->tempfile);
> +		preq->packfile = NULL;
> +	} else if (preq->packfile) {
>  		fclose(preq->packfile);
>  		preq->packfile = NULL;
>  	}

OK. I think this is correct, though see my comments elsewhere in the
thread.

> @@ -2688,7 +2691,10 @@ int finish_http_pack_request(struct http_pack_request *preq)
>  	int tmpfile_fd;
>  	int ret = 0;
>  
> -	fclose(preq->packfile);
> +	if (preq->tempfile)
> +		close_tempfile_gently(preq->tempfile);
> +	else
> +		fclose(preq->packfile);
>  	preq->packfile = NULL;

OK, and this is correct because preq->packfile is just an alias for
preq->tempfile.fp when the tempfile is valid. The NULL assignment is
important here so that the release() function doesn't double-free.

> -struct http_pack_request *new_http_pack_request(
> -	const unsigned char *packed_git_hash, const char *base_url) {
> -
> -	struct strbuf buf = STRBUF_INIT;
> -
> -	end_url_with_slash(&buf, base_url);
> -	strbuf_addf(&buf, "objects/pack/pack-%s.pack",
> -		hash_to_hex(packed_git_hash));
> -	return new_direct_http_pack_request(packed_git_hash,
> -					    strbuf_detach(&buf, NULL));
> -}

This hunk puzzled me at first, but it's because we used to just be a
wrapper for the "direct" variant, and now the two will share a single
static helper. That might have been a little more clear as a preparatory
patch, but OK.

> +	if (resumable) {
> +		odb_pack_name(the_repository, &preq->tmpfile,
> +			      packed_git_hash, "pack");
> +		strbuf_addstr(&preq->tmpfile, ".temp");
> +		preq->packfile = fopen(preq->tmpfile.buf, "a");
> +	} else {
> +		strbuf_addf(&preq->tmpfile, "%s/pack/tmp_pack_XXXXXX",
> +			    repo_get_object_directory(the_repository));
> +		preq->tempfile = mks_tempfile_m(preq->tmpfile.buf, 0444);
> +		if (preq->tempfile) {
> +			strbuf_reset(&preq->tmpfile);
> +			strbuf_addstr(&preq->tmpfile,
> +				      get_tempfile_path(preq->tempfile));
> +			preq->packfile = fdopen_tempfile(preq->tempfile, "w");
> +		}
> +	}
>  	if (!preq->packfile) {
>  		error("Unable to open local file %s for pack",
>  		      preq->tmpfile.buf);

OK, and this is the meat of the change. We usually use odb_mkstemp() for
tmp_pack_* files, but that annoyingly doesn't give you a tempfile
struct. So setting up your own filename and using mks_tempfile_m() makes
sense here.

The error path is a little funny, but we catch it in the context when
preq->packfile is NULL. Good.

> @@ -2766,8 +2776,9 @@ struct http_pack_request *new_direct_http_pack_request(
>  	 * If there is data present from a previous transfer attempt,
>  	 * resume where it left off
>  	 */
> -	prev_posn = ftello(preq->packfile);
> -	if (prev_posn>0) {
> +	if (resumable)
> +		prev_posn = ftello(preq->packfile);
> +	if (prev_posn > 0) {

I think this is not technically necessary, as ftello() would just return
"0" for our newly-created file. But it does make the intent clear.

> @@ -2779,12 +2790,28 @@ struct http_pack_request *new_direct_http_pack_request(
>  	return preq;
>  
>  abort:
> -	strbuf_release(&preq->tmpfile);
> -	free(preq->url);
> -	free(preq);
> +	release_http_pack_request(preq);
>  	return NULL;
>  }

OK, now we have potentially more to free, so we rely on the release
function. That could cause problems if we jump to this abort label when
the struct isn't fully initialized. I think it is OK, though. We zero
the whole thing, so the extra fields that the release() function
considers will just be ignored.

> diff --git a/http.h b/http.h
> index 729c51904d..2c900779f5 100644
> --- a/http.h
> +++ b/http.h
> @@ -224,6 +224,7 @@ struct http_pack_request {
>  
>  	FILE *packfile;
>  	struct strbuf tmpfile;
> +	struct tempfile *tempfile;
>  	struct active_request_slot *slot;
>  	struct curl_slist *headers;

Yuck, now we have "tempfile" and "tmpfile" with two different types and
totally different semantics (and even when "tempfile" is in use,
"tmpfile" is still meaningful!).

Can we even just call the second one non_resumable_tempfile or
something? It's a mouthful, but it makes it less likely to confuse the
two.

> +	# Hold the first download before it is indexed, so that the second
> +	# download installs the pack first.
> +	{
> +		(
> +			if ! PATH="$TRASH_DIRECTORY:$PATH" \
> +			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
> +			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
> +			git -C packfileclient-concurrent http-fetch \
> +				--packfile="$packhash" \
> +				--index-pack-arg=wait-index-pack \
> +				--index-pack-arg=--stdin \
> +				--index-pack-arg=--keep \
> +				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
> +			then
> +				echo failed >"$TRASH_DIRECTORY/first-ready" &&
> +				exit 1
> +			fi
> +		) &
> +		first_pid=$!
> +	} &&

OK. I wonder if it would be simpler and a more robust test if rather
than writing the correct bytes (and then waiting), the first process
just wrote total garbage. Then we'd be sure the other process is not
reading it, because it would definitely corrupt their input.

I dunno. This is a more realistic scenario, so in that sense maybe it is
more interesting.

> +	test_when_finished "
> +		echo continue >&9
> +		wait $first_pid 2>/dev/null || :
> +		exec 8>&-
> +		exec 9>&-
> +		rm -f first-ready first-continue git-wait-index-pack
> +	" &&
> [...]

The rest of the fifo handling looks plausibly correct. This is a tricky
area and it's common to introduce funky races, but I didn't see anything
wrong, and it passed a few dozen rounds of --stress.

> @@ -313,7 +381,9 @@ test_expect_success 'http-fetch --packfile with corrupt pack' '
>  	git init packfileclient &&
>  	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git && ls objects/pack/pack-*.pack) &&
>  	test_must_fail git -C packfileclient http-fetch --packfile \
> -		"$HTTPD_URL"/dumb/repo_bad1.git/$p
> +		"$HTTPD_URL"/dumb/repo_bad1.git/$p &&
> +	find packfileclient/.git/objects/pack -name "tmp_pack_*" -print >tmpfiles &&
> +	test_must_be_empty tmpfiles
>  '

OK, so here we just detect that we cleaned up after ourselves. Makes
sense.

-Peff
Jeff KingJul 14, 2026, 07:12 UTC in reply to Ted Nyman on lore

Re: [PATCH 2/2] fetch-pack: accept "pack" output for packfile URIs

On Mon, Jul 13, 2026 at 03:34:43PM -0700, Ted Nyman wrote:
Show 14 quoted lines
> When "index-pack --keep" creates a .keep file, it reports
> "keep<TAB><hash>". If the file already exists, index-pack leaves it
> untouched and reports "pack<TAB><hash>" instead.
> 
> Since dd4b732df7 (upload-pack: send part of packfile response as uri,
> 2020-06-10), fetch-pack has accepted only the "keep" form for packs
> downloaded through packfile URIs. A concurrent fetch can install the
> same pack and create its .keep file before another process reaches
> index-pack. The latter process then fails even though index-pack
> completed successfully.
> 
> Accept both successful forms. Add a path to pack_lockfiles only for the
> "keep" form, so cleanup removes only a keep file created by the current
> process and preserves a pre-existing one.
OK, that all makes sense.
Show 8 quoted lines
>  	for (i = 0; i < packfile_uris.nr; i++) {
> +		int created_keep = 0;
>  		int j;
>  		struct child_process cmd = CHILD_PROCESS_INIT;
> -		char packname[GIT_MAX_HEXSZ + 1];
> +		char packname[GIT_MAX_HEXSZ + 6];
> +		const char *packhash;
> +		const int packname_len = the_hash_algo->hexsz + 6;

The "+ 6" here is gross, but not really any more than the bare "5" in the original code.

Calling it "packhash" made me wonder about this line of code:
> -		if (memcmp(packfile_uris.items[i].string, packname,
> +		if (memcmp(packfile_uris.items[i].string, packhash,

Surely we need to change more than this if we now have the hash rather than the whole packname? But no, the original code was really just storing the hash in packname.

Which was rather misleading, but it is not much better after your patch. Now packname still just has the packhash, along with the extra keep/pack marker.

Would a more generic name like "cmd_output" or something make sense? I also think this would all be much nicer with a strbuf (which would let us get rid of the magic numbers), but that is a slightly larger refactor:

Show changes to fetch-pack.c +5 −10
diff --git a/fetch-pack.c b/fetch-pack.c
index 1e8461d07e..5f94f35c30 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1890,9 +1890,8 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		int created_keep = 0;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 6];
+		struct strbuf cmd_output = STRBUF_INIT;
 		const char *packhash;
-		const int packname_len = the_hash_algo->hexsz + 6;
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1910,14 +1909,11 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, packname_len) != packname_len ||
-		    packname[packname_len - 1] != '\n')
-			die("fetch-pack: expected pack or keep, TAB, hash, "
-			    "then LF in http-fetch output");
-		packname[packname_len - 1] = '\0';
-		if (skip_prefix(packname, "keep\t", &packhash))
+		if (strbuf_read(&cmd_output, cmd.out, 0) < 0)
+			die("failed to read http-fetch output");
+		if (skip_prefix(cmd_output.buf, "keep\t", &packhash))
 			created_keep = 1;
-		else if (!skip_prefix(packname, "pack\t", &packhash))
+		else if (!skip_prefix(cmd_output.buf, "pack\t", &packhash))
 			die("fetch-pack: expected pack or keep, TAB, hash, "
 			    "then LF in http-fetch output");
 


BTW, two things that puzzled me while poking at your patch (but neither
I think are new or the fault of your patch):

  1. git-http-fetch documents --index-pack-args, but the actual option
     is the singular --index-pack-arg. The caller in fetch-pack
     obviously uses the one that works.

  2. The code you're touching insists on reading "keep" in the output,
     but wouldn't that depend on feeding "--keep" via index_pack_args? I
     didn't see immediately where we set it, but I certainly don't think
     your patch could be making anything worse here (since the existing
     code would have just died upon seeing a "pack" line). I think it
     happens as a side effect in get_pack(), which is...subtle. But
     again, not anything new.

-Peff
Jeff KingJul 14, 2026, 07:13 UTC in reply to Jeff King on lore

Re: [PATCH 2/2] fetch-pack: accept "pack" output for packfile URIs

On Tue, Jul 14, 2026 at 03:12:31AM -0400, Jeff King wrote:
Show 6 quoted lines
> Would a more generic name like "cmd_output" or something make sense? I
> also think this would all be much nicer with a strbuf (which would let
> us get rid of the magic numbers), but that is a slightly larger
> refactor:
> 
> diff --git a/fetch-pack.c b/fetch-pack.c
In case anybody does pursue this, it is obviously missing this bit:
Show changes to fetch-pack.c +2 −1
diff --git a/fetch-pack.c b/fetch-pack.c
index 5f94f35c30..359740f231 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1935,6 +1935,8 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 						 xstrfmt("%s/pack/pack-%s.keep",
 							 repo_get_object_directory(the_repository),
 							 packhash));
+
+		strbuf_release(&cmd_output);
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);

to avoid a leak.

-Peff
Junio C HamanoJul 14, 2026, 18:10 UTC in reply to Jeff King on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

Jeff King <peff@peff.net> writes:
Show 44 quoted lines
> On Mon, Jul 13, 2026 at 06:58:24PM -0700, Ted Nyman wrote:
>
>> > Are there better ways for these processes to coordinate with each
>> > other? Instead of appending to the file, what if the second process
>> > uses a predictable temporary name (which we already use) to open a
>> > new file with O_CREAT | O_EXCL to avoid this redundant work?
>> 
>> Using the existing pack-<hash>.pack.temp name with O_CREAT | O_EXCL
>> would prevent concurrent writes, but EEXIST alone would not
>> distinguish an in-progress download from one left by an earlier
>> failed or interrupted invocation. The existing .pack.temp name is not
>> covered by the tmp_* pruning path, so simply waiting for it to
>> disappear could leave a fetch stuck after a crash.
>
> A few thoughts:
>
>   - Using O_EXCL makes this essentially a lockfile. So we could apply
>     the logic used elsewhere for lockfiles, like auto-removing files
>     with ancient mtimes. Or we could even go all-in with a pid check for
>     liveness; most of Git's lockfiles don't do that, but at least one
>     does (the background auto-gc lock).
>
>   - If we're not already using a name which is auto-cleaned during
>     maintenance, we probably ought to be. Leaving aside concurrency
>     issues, nobody would ever clean up the on-disk cruft.
>
>     But of course the original code here is intentionally _not_ using a
>     name we'd clean up, because it wants to be able to resume an
>     interrupted transfer.  And you're explicitly breaking that for the
>     packfile URI case.
>
>     Is that a cost we're OK with paying? Fixing it opens up that same
>     coordination can of worms. You have to tell the difference a
>     concurrent writer and a previous dead one (whose work you can
>     resume).
>
>     It does feel weird that we'd do one thing for dumb-http and another
>     for packfile URIs. Wouldn't they suffer from the same concurrency
>     and resumption problems?
> ...
>
> If we're OK with killing the ability to resume, then yeah, I think it
> would make sense to start simple and un-break things. And then put a
> coordination layer on top later (or never if nobody cares enough).

I share that sentiment. I am not entirely convinced by Ted's response, since a major goal of the packfile URI feature, as I understand it, is to allow the use of resumable protocols for large transfers. The proposed change deliberately closes the door on resuming interrupted transfers, whether manually or, with additional code in the future, automatically.

Ted NymanJul 14, 2026, 18:31 UTC in reply to Junio C Hamano on lore

Re: [PATCH 1/2] http: use unique tempfiles for packfile URI downloads

> I share that sentiment. I am not entirely convinced by Ted's
> response, since a major goal of the packfile URI feature, as I
> understand it, is to allow the use of resumable protocols for
> large transfers.
Agreed. I was too quick to dismiss the loss of resumption.

I'll take another look at preserving the predictable partial pack while preventing concurrent writers, including the handoff and stale-file cases Peff raised. Dumb HTTP has the same underlying concurrency issue, so I'll keep that path in mind as well before sending a reroll.

Thanks, Ted

Ted NymanJul 14, 2026, 18:38 UTC in reply to Jeff King on lore

Re: [PATCH 2/2] fetch-pack: accept "pack" output for packfile URIs

> I also think this would all be much nicer with a strbuf (which would
> let us get rid of the magic numbers), but that is a slightly larger
> refactor:

Using a strbuf makes sense. One wrinkle, I think, is that with transfer.fsckobjects enabled, index-pack can emit dangling .gitmodules OIDs after the initial pack/keep line, which parse_gitmodules_oids() still needs to read from cmd.out. Would strbuf_getwholeline_fd() be a better fit here, so we don't consume those with strbuf_read()?

I'll also fix the --index-pack-args documentation while rerolling.

Thanks, Ted

Jeff KingJul 14, 2026, 21:47 UTC in reply to Ted Nyman on lore

Re: [PATCH 2/2] fetch-pack: accept "pack" output for packfile URIs

On Tue, Jul 14, 2026 at 11:38:56AM -0700, Ted Nyman wrote:
Show 9 quoted lines
> > I also think this would all be much nicer with a strbuf (which would
> > let us get rid of the magic numbers), but that is a slightly larger
> > refactor:
> 
> Using a strbuf makes sense. One wrinkle, I think, is that with
> transfer.fsckobjects enabled, index-pack can emit dangling .gitmodules
> OIDs after the initial pack/keep line, which parse_gitmodules_oids()
> still needs to read from cmd.out. Would strbuf_getwholeline_fd() be a
> better fit here, so we don't consume those with strbuf_read()?

Ah, yeah, I didn't think about whether it might have more output. I _think_ it actually works just fine with more output because the memcmp() is limited to the hash algo's hex_sz. For the same reason what I posted works even though it has the trailing newline.

It is a bit subtle, though. Using getwholeline_fd would work (though you still have the trailing newline subtlety). Or maybe just using strbuf_setlen() to cut off the output (ironically it is probably more efficient to read the whole thing in and then chomp it, since getwholeline_fd will read() one char at a time).

The "cleanest" thing is perhaps xfdopen() followed by strbuf_getline(), but maybe that's overkill.

I'd be happy with any of the solutions. Or even just keeping the magic numbers but maybe with a comment explaining what the heck "6" means.

> I'll also fix the --index-pack-args documentation while rerolling.
Great, thanks.
-Peff
Ted NymanJul 20, 2026, 22:33 UTC in reply to Ted Nyman on lore

[PATCH v2 0/2] packfile URIs: support concurrent downloads

Packfile URI and dumb HTTP downloads stage packs at objects/pack/pack-<hash>.pack.temp so an interrupted transfer can resume. Two Git processes fetching the same pack into one object database can append to that file concurrently, which can corrupt the temporary pack.

Following Peff's suggestion in the v1 discussion, the first patch keeps the predictable temporary name but opens the file without append mode. Each downloader keeps its own file offset, so overlapping responses for the same pack write the same bytes at the same offsets. This preserves resumable downloads and protects both packfile URI and ordinary dumb HTTP requests. Concurrent downloaders may still transfer the same suffix; this series avoids adding cross-process coordination and a stale-owner policy.

A second downloader can also find that the partial pack has completed and request a range starting at EOF. Servers may respond with HTTP 416 in that case. Treat that response as a completed download and let index-pack validate the pack. Keep the open descriptor for indexing so another downloader can safely remove the temporary path.

The second patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process. Read only the prefix and hash so any following fsck output remains available to fetch-pack.

The tests cover resumption, a completed partial returning 416, overlapping 200 and 206 responses, and a pre-existing .keep file.

Changes since v1:
  * Preserve resumability by removing append mode instead of using a
    unique temporary file for each download.
  * Handle the EOF-range/416 and concurrent-unlink cases, including the
    Windows sharing behavior.
  * Add a deterministic overlapping-download regression test.
  * Read the pack/keep prefix and hash without consuming later fsck
    output, as suggested during review.
  * Correct the stale --index-pack-args documentation and error text;
    the repeatable --index-pack-arg option is already supported.
The v1 discussion is at:
  https://lore.kernel.org/git/cover.1783982021.git.tnyman@openai.com/
Ted Nyman (2):
  http: avoid concurrent appends to partial packs
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  13 +-
 fetch-pack.c                      |  33 +++--
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  53 ++++---
 t/t5550-http-fetch-dumb.sh        | 223 ++++++++++++++++++++++++++++++
 t/t5702-protocol-v2.sh            |  31 +++++
 8 files changed, 320 insertions(+), 46 deletions(-)
Range-diff against v1:
1:  32eb9b0831 ! 1:  160a9b9fd0 http: use unique tempfiles for packfile URI downloads
    @@ Metadata
     Author: Ted Nyman <tnyman@openai.com>
     
      ## Commit message ##
    -    http: use unique tempfiles for packfile URI downloads
    +    http: avoid concurrent appends to partial packs
     
    -    Since 8d5d2a34df (http-fetch: support fetching packfiles by URL,
    -    2020-06-10), packfile URI downloads have been staged at
    -    objects/pack/pack-<hash>.pack.temp.
    +    Pack requests stage downloads in a predictable partial-pack file so an
    +    interrupted transfer can be resumed. Both packfile URI and ordinary dumb
    +    HTTP requests use this staging path. Opening it in append mode lets
    +    concurrent fetches interleave their writes, corrupting the pack or
    +    causing a later fetch to request a range at EOF.
     
    -    The path is derived from the advertised pack hash. Two processes
    -    fetching the same pack into a shared object database therefore open the
    -    same file for append. Their writes can corrupt the temporary pack. If
    -    one process arrives after the other has completed the download, it may
    -    instead try to resume at EOF, which some HTTP servers reject with 416.
    +    Open the partial pack read-write, seek to its current end, and retain a
    +    per-descriptor offset for incoming data. Reopen newly created partial
    +    packs without O_CREAT so Windows permits concurrent unlink, and keep the
    +    descriptor for index-pack when another downloader removes the staging
    +    path. Accept HTTP 416 when a partial pack is already complete.
     
    -    Use the tempfile API to give direct packfile URI downloads unique
    -    temporary files. Keep the deterministic path for ordinary dumb HTTP
    -    pack requests, which use it to resume a partial download left by an
    -    earlier invocation.
    -
    -    This means that a packfile URI download cannot be resumed by a later
    -    invocation. A retry starts with an empty temporary file instead.
    -
    -    Add a test which pauses one process after downloading the pack and
    -    starts another process using the same object database.
    +    Exercise resumed transfers, EOF ranges, and overlapping 200 and 206
    +    responses. Clarify the staging-key documentation and correct the stale
    +    --index-pack-args spelling in the documentation and error messages; the
    +    repeatable --index-pack-arg option is already accepted.
     
         Signed-off-by: Ted Nyman <tnyman@openai.com>
    -    Signed-off-by: Junio C Hamano <gitster@pobox.com>
     
      ## Documentation/git-http-fetch.adoc ##
     @@ Documentation/git-http-fetch.adoc: commit-id::
    @@ Documentation/git-http-fetch.adoc: commit-id::
     -	The hash is used to determine the name of the temporary file and is
     -	arbitrary. The output of index-pack is printed to stdout. Requires
     -	--index-pack-args.
    -+	The hash is arbitrary. The output of index-pack is printed to stdout.
    -+	Requires --index-pack-args.
    ++	The hash is used to determine the name of the temporary file. It need
    ++	not be the pack hash, but it must uniquely identify the pack contents
    ++	for resumption. The output of index-pack is printed to stdout. Requires
    ++	one or more --index-pack-arg options.
    + 
    +---index-pack-args=<args>::
    +-	For internal use only. The command to run on the contents of the
    +-	downloaded pack. Arguments are URL-encoded separated by spaces.
    ++--index-pack-arg=<arg>::
    ++	For internal use only. An argument to the command run on the contents
    ++	of the downloaded pack. This option can be specified multiple times.
      
    - --index-pack-args=<args>::
    - 	For internal use only. The command to run on the contents of the
    + --recover::
    + 	Verify that everything reachable from target is fetched.  Used after
     
    - ## http.c ##
    -@@ http.c: int http_get_info_packs(const char *base_url, struct packfile_list *packs)
    + ## http-fetch.c ##
    +@@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,
      
    - void release_http_pack_request(struct http_pack_request *preq)
    - {
    --	if (preq->packfile) {
    -+	if (preq->tempfile) {
    -+		delete_tempfile(&preq->tempfile);
    -+		preq->packfile = NULL;
    -+	} else if (preq->packfile) {
    - 		fclose(preq->packfile);
    - 		preq->packfile = NULL;
    + 	if (start_active_slot(preq->slot)) {
    + 		run_active_slot(preq->slot);
    +-		if (results.curl_result != CURLE_OK) {
    ++		if (results.curl_result != CURLE_OK &&
    ++		    results.http_code != 416) {
    + 			struct url_info url;
    + 			char *nurl = url_normalize(preq->url, &url);
    + 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
    +@@ http-fetch.c: int cmd_main(int argc, const char **argv)
    + 
    + 	if (packfile) {
    + 		if (!index_pack_args.nr)
    +-			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
    ++			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
    + 
    + 		fetch_single_packfile(&packfile_hash, argv[arg],
    + 				      index_pack_args.v);
    +@@ http-fetch.c: int cmd_main(int argc, const char **argv)
      	}
    + 
    + 	if (index_pack_args.nr)
    +-		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
    ++		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
    + 
    + 	if (commits_on_stdin) {
    + 		commits = walker_targets_stdin(&commit_id, &write_ref);
    +
    + ## http-push.c ##
    +@@ http-push.c: static void finish_request(struct transfer_request *request)
    + 
    + 	} else if (request->state == RUN_FETCH_PACKED) {
    + 		int fail = 1;
    +-		if (request->curl_result != CURLE_OK) {
    ++		if (request->curl_result != CURLE_OK &&
    ++		    request->http_code != 416) {
    + 			fprintf(stderr, "Unable to get pack file %s\n%s",
    + 				request->url, curl_errorstr);
    + 		} else {
    +
    + ## http-walker.c ##
    +@@ http-walker.c: static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
    + 
    + 	if (start_active_slot(preq->slot)) {
    + 		run_active_slot(preq->slot);
    +-		if (results.curl_result != CURLE_OK) {
    ++		if (results.curl_result != CURLE_OK &&
    ++		    results.http_code != 416) {
    + 			error("Unable to get pack file %s\n%s", preq->url,
    + 			      curl_errorstr);
    + 			goto abort;
    +
    + ## http.c ##
     @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)
      	int tmpfile_fd;
      	int ret = 0;
      
    --	fclose(preq->packfile);
    -+	if (preq->tempfile)
    -+		close_tempfile_gently(preq->tempfile);
    -+	else
    -+		fclose(preq->packfile);
    ++	/* Another downloader may unlink the staging path while we index it. */
    ++	tmpfile_fd = xdup(fileno(preq->packfile));
    + 	fclose(preq->packfile);
      	preq->packfile = NULL;
    +-
    +-	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
    ++	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
    ++		die_errno("unable to seek local file %s for pack",
    ++			  preq->tmpfile.buf);
      
    - 	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
    + 	ip.git_cmd = 1;
    + 	ip.in = tmpfile_fd;
     @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)
    + 	else
    + 		ip.no_stdout = 1;
      
    - cleanup:
    - 	close(tmpfile_fd);
    --	unlink(preq->tmpfile.buf);
    -+	if (preq->tempfile)
    -+		delete_tempfile(&preq->tempfile);
    -+	else
    -+		unlink(preq->tmpfile.buf);
    +-	if (run_command(&ip)) {
    ++	if (run_command(&ip))
    + 		ret = -1;
    +-		goto cleanup;
    +-	}
    +-
    +-cleanup:
    +-	close(tmpfile_fd);
    + 	unlink(preq->tmpfile.buf);
      	return ret;
      }
    - 
    -@@ http.c: void http_install_packfile(struct packed_git *p,
    - 	packfile_store_add_pack(files->packed, p);
    - }
    - 
    --struct http_pack_request *new_http_pack_request(
    --	const unsigned char *packed_git_hash, const char *base_url) {
    --
    --	struct strbuf buf = STRBUF_INIT;
    --
    --	end_url_with_slash(&buf, base_url);
    --	strbuf_addf(&buf, "objects/pack/pack-%s.pack",
    --		hash_to_hex(packed_git_hash));
    --	return new_direct_http_pack_request(packed_git_hash,
    --					    strbuf_detach(&buf, NULL));
    --}
    --
    --struct http_pack_request *new_direct_http_pack_request(
    --	const unsigned char *packed_git_hash, char *url)
    -+static struct http_pack_request *new_http_pack_request_for_url(
    -+	const unsigned char *packed_git_hash, char *url, int resumable)
    +@@ http.c: struct http_pack_request *new_http_pack_request(
    + struct http_pack_request *new_direct_http_pack_request(
    + 	const unsigned char *packed_git_hash, char *url)
      {
    - 	off_t prev_posn = 0;
    +-	off_t prev_posn = 0;
    ++	off_t prev_posn;
      	struct http_pack_request *preq;
    -@@ http.c: struct http_pack_request *new_direct_http_pack_request(
    ++	int fd;
      
    + 	CALLOC_ARRAY(preq, 1);
    + 	strbuf_init(&preq->tmpfile, 0);
    +-
      	preq->url = url;
      
    --	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
    --	strbuf_addstr(&preq->tmpfile, ".temp");
    + 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
    + 	strbuf_addstr(&preq->tmpfile, ".temp");
     -	preq->packfile = fopen(preq->tmpfile.buf, "a");
    -+	if (resumable) {
    -+		odb_pack_name(the_repository, &preq->tmpfile,
    -+			      packed_git_hash, "pack");
    -+		strbuf_addstr(&preq->tmpfile, ".temp");
    -+		preq->packfile = fopen(preq->tmpfile.buf, "a");
    -+	} else {
    -+		strbuf_addf(&preq->tmpfile, "%s/pack/tmp_pack_XXXXXX",
    -+			    repo_get_object_directory(the_repository));
    -+		preq->tempfile = mks_tempfile_m(preq->tmpfile.buf, 0444);
    -+		if (preq->tempfile) {
    -+			strbuf_reset(&preq->tmpfile);
    -+			strbuf_addstr(&preq->tmpfile,
    -+				      get_tempfile_path(preq->tempfile));
    -+			preq->packfile = fdopen_tempfile(preq->tempfile, "w");
    +-	if (!preq->packfile) {
    +-		error("Unable to open local file %s for pack",
    +-		      preq->tmpfile.buf);
    ++	/* Reopen without O_CREAT so MinGW permits another writer to unlink it. */
    ++	for (;;) {
    ++		fd = open(preq->tmpfile.buf, O_RDWR);
    ++		if (fd >= 0 || errno != ENOENT)
    ++			break;
    ++		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
    ++		if (fd >= 0) {
    ++			close(fd);
    ++			continue;
     +		}
    ++		if (errno != EEXIST)
    ++			break;
    ++	}
    ++	if (fd < 0) {
    ++		error_errno("unable to open local file %s for pack",
    ++			    preq->tmpfile.buf);
    ++		goto abort;
     +	}
    - 	if (!preq->packfile) {
    - 		error("Unable to open local file %s for pack",
    - 		      preq->tmpfile.buf);
    ++	prev_posn = lseek(fd, 0, SEEK_END);
    ++	if (prev_posn < 0) {
    ++		error_errno("unable to seek local file %s for pack",
    ++			    preq->tmpfile.buf);
    ++		close(fd);
    + 		goto abort;
    + 	}
    ++	preq->packfile = xfdopen(fd, "w");
    + 
    + 	preq->slot = get_active_slot();
    + 	preq->headers = object_request_headers();
     @@ http.c: struct http_pack_request *new_direct_http_pack_request(
    - 	 * If there is data present from a previous transfer attempt,
    - 	 * resume where it left off
    - 	 */
    + 	curl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);
    + 	curl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);
    + 
    +-	/*
    +-	 * If there is data present from a previous transfer attempt,
    +-	 * resume where it left off
    +-	 */
     -	prev_posn = ftello(preq->packfile);
     -	if (prev_posn>0) {
    -+	if (resumable)
    -+		prev_posn = ftello(preq->packfile);
     +	if (prev_posn > 0) {
      		if (http_is_verbose)
      			fprintf(stderr,
      				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
    -@@ http.c: struct http_pack_request *new_direct_http_pack_request(
    - 	return preq;
    - 
    - abort:
    --	strbuf_release(&preq->tmpfile);
    --	free(preq->url);
    --	free(preq);
    -+	release_http_pack_request(preq);
    - 	return NULL;
    - }
    - 
    -+struct http_pack_request *new_http_pack_request(
    -+	const unsigned char *packed_git_hash, const char *base_url)
    -+{
    -+	struct strbuf buf = STRBUF_INIT;
    -+
    -+	end_url_with_slash(&buf, base_url);
    -+	strbuf_addf(&buf, "objects/pack/pack-%s.pack",
    -+		hash_to_hex(packed_git_hash));
    -+	return new_http_pack_request_for_url(packed_git_hash,
    -+					     strbuf_detach(&buf, NULL), 1);
    -+}
    -+
    -+struct http_pack_request *new_direct_http_pack_request(
    -+	const unsigned char *packed_git_hash, char *url)
    -+{
    -+	return new_http_pack_request_for_url(packed_git_hash, url, 0);
    -+}
    -+
    - /* Helpers for fetching objects (loose) */
    - static size_t fwrite_sha1_file(char *ptr, size_t eltsize, size_t nmemb,
    - 			       void *data)
    -
    - ## http.h ##
    -@@ http.h: struct http_pack_request {
    - 
    - 	FILE *packfile;
    - 	struct strbuf tmpfile;
    -+	struct tempfile *tempfile;
    - 	struct active_request_slot *slot;
    - 	struct curl_slist *headers;
    - };
     
      ## t/t5550-http-fetch-dumb.sh ##
     @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
      	git -C packfileclient cat-file -e "$HASH"
      '
      
    -+test_expect_success PIPE 'concurrent http-fetch --packfile' '
    ++test_expect_success 'http-fetch --packfile resumes a partial download' '
    ++	git init packfileclient-resume &&
    ++	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
    ++		ls objects/pack/pack-*.pack) &&
    ++	tmpfile="packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
    ++	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
    ++	GIT_TRACE_CURL="$TRASH_DIRECTORY/resume.trace" \
    ++	git -C packfileclient-resume http-fetch --packfile="$ARBITRARY" \
    ++		--index-pack-arg=index-pack --index-pack-arg=--stdin \
    ++		--index-pack-arg=--keep \
    ++		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
    ++	test_grep "Range: bytes=64-" resume.trace &&
    ++	test_path_is_missing "$tmpfile" &&
    ++	git -C packfileclient-resume cat-file -e "$HASH"
    ++'
    ++
    ++test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
     +	git init packfileclient-concurrent &&
    -+	HASH=$(git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git rev-parse HEAD) &&
     +	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
     +		ls objects/pack/pack-*.pack) &&
     +	packhash=$(basename "$p" .pack) &&
     +	packhash=${packhash#pack-} &&
    -+
    ++	tmpfile="packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
    ++	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
     +	mkfifo first-ready first-continue &&
     +	exec 8<>first-ready &&
     +	exec 9<>first-continue &&
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
     +	exec git index-pack "$@"
     +	EOF
    -+
    -+	# Hold the first download before it is indexed, so that the second
    -+	# download installs the pack first.
     +	{
     +		(
     +			if ! PATH="$TRASH_DIRECTORY:$PATH" \
     +			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
     +			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
    -+			git -C packfileclient-concurrent http-fetch \
    -+				--packfile="$packhash" \
    ++			GIT_TRACE_CURL="$TRASH_DIRECTORY/first.trace" \
    ++			git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
     +				--index-pack-arg=wait-index-pack \
    -+				--index-pack-arg=--stdin \
    -+				--index-pack-arg=--keep \
    ++				--index-pack-arg=--stdin --index-pack-arg=--keep \
     +				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
     +			then
     +				echo failed >"$TRASH_DIRECTORY/first-ready" &&
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	} &&
     +	test_when_finished "
     +		echo continue >&9
    ++		kill $first_pid 2>/dev/null || :
     +		wait $first_pid 2>/dev/null || :
     +		exec 8>&-
     +		exec 9>&-
     +		rm -f first-ready first-continue git-wait-index-pack
     +	" &&
    -+
     +	read ready <&8 &&
     +	test "$ready" = ready &&
    -+	git -C packfileclient-concurrent http-fetch \
    -+		--packfile="$packhash" \
    ++	GIT_TRACE_CURL="$TRASH_DIRECTORY/second.trace" \
    ++	git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
     +		--index-pack-arg=index-pack \
    -+		--index-pack-arg=--stdin \
    -+		--index-pack-arg=--keep \
    ++		--index-pack-arg=--stdin --index-pack-arg=--keep \
     +		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
     +	echo continue >&9 &&
     +	wait "$first_pid" &&
    -+
     +	printf "pack\t%s\n" "$packhash" >expect &&
     +	test_cmp expect first.out &&
     +	printf "keep\t%s\n" "$packhash" >expect &&
     +	test_cmp expect second.out &&
    -+	test_path_is_missing \
    -+		"packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
    -+	find packfileclient-concurrent/.git/objects/pack \
    -+		-name "tmp_pack_*" -print >tmpfiles &&
    -+	test_must_be_empty tmpfiles &&
    ++	test_grep "Range: bytes=64-" first.trace &&
    ++	test_grep "Range: bytes=[0-9]*-" second.trace &&
    ++	test_grep "HTTP/[0-9.]* 416" second.trace &&
    ++	test_path_is_missing "$tmpfile" &&
     +	git -C packfileclient-concurrent cat-file -e "$HASH"
     +'
    ++
    ++test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
    ++	git init packfileclient-overlap &&
    ++	blob=$(test-tool genrandom pack-overlap 2m |
    ++		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
    ++			hash-object -w --stdin) &&
    ++	packhash=$(printf "%s\n" "$blob" |
    ++		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
    ++			pack-objects "$TRASH_DIRECTORY/overlap-pack") &&
    ++	pack="$TRASH_DIRECTORY/overlap-pack-$packhash.pack" &&
    ++	tmpfile="packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp" &&
    ++	mkfifo server-ready first-ready &&
    ++	exec 7<>server-ready &&
    ++	exec 8<>first-ready &&
    ++	write_script slow-pack-server "$PERL_PATH" <<-\EOF &&
    ++	use strict;
    ++	use warnings;
    ++	use IO::Socket::INET;
    ++
    ++	my ($packfile, $server_ready, $first_ready) = @ARGV;
    ++	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
    ++	my $pack = do { local $/; <$in> };
    ++	close($in) or die "close $packfile: $!";
    ++	my $server = IO::Socket::INET->new(LocalAddr => "127.0.0.1",
    ++		LocalPort => 0, Proto => "tcp", Listen => 2, ReuseAddr => 1)
    ++		or die "listen: $!";
    ++
    ++	sub signal_ready {
    ++		my ($file, $value) = @_;
    ++		open(my $out, ">", $file) or die "open $file: $!";
    ++		print $out "$value\n" or die "write $file: $!";
    ++		close($out) or die "close $file: $!";
    ++	}
    ++
    ++	sub write_all {
    ++		my ($out, $data) = @_;
    ++		my $offset = 0;
    ++		while ($offset < length($data)) {
    ++			my $written = syswrite($out, $data,
    ++				length($data) - $offset, $offset);
    ++			defined($written) && $written or die "write response: $!";
    ++			$offset += $written;
    ++		}
    ++	}
    ++
    ++	sub start_response {
    ++		my $out = $server->accept() or die "accept: $!";
    ++		<$out> or die "read request: $!";
    ++		my $start = 0;
    ++		while (<$out>) {
    ++			last if /^\r?\n$/;
    ++			$start = $1 if /^Range: bytes=(\d+)-/i;
    ++		}
    ++		$start < length($pack) or die "invalid range $start";
    ++		my $length = length($pack) - $start;
    ++		my $middle = int($length / 2);
    ++		my $status = $start ? "206 Partial Content" : "200 OK";
    ++		my $headers = "HTTP/1.1 $status\r\n" .
    ++			"Content-Length: $length\r\n" .
    ++			($start ? "Content-Range: bytes $start-" .
    ++				(length($pack) - 1) . "/" . length($pack) . "\r\n" : "") .
    ++			"Connection: close\r\n\r\n";
    ++		write_all($out, $headers);
    ++		write_all($out, substr($pack, $start, $middle));
    ++		return ($out, $start + $middle);
    ++	}
    ++
    ++	signal_ready($server_ready, $server->sockport());
    ++	my ($first, $first_pos) = start_response();
    ++	signal_ready($first_ready, "ready");
    ++	my ($second, $second_pos) = start_response();
    ++	write_all($first, substr($pack, $first_pos));
    ++	write_all($second, substr($pack, $second_pos));
    ++	close($first) or die "close first response: $!";
    ++	close($second) or die "close second response: $!";
    ++	EOF
    ++	{
    ++		(
    ++			if ! "$TRASH_DIRECTORY/slow-pack-server" "$pack" \
    ++				"$TRASH_DIRECTORY/server-ready" \
    ++				"$TRASH_DIRECTORY/first-ready"
    ++			then
    ++				echo failed >"$TRASH_DIRECTORY/server-ready" &&
    ++				echo failed >"$TRASH_DIRECTORY/first-ready" &&
    ++				exit 1
    ++			fi
    ++		) >server.log 2>&1 &
    ++		server_pid=$!
    ++	} &&
    ++	test_when_finished "
    ++		kill $server_pid 2>/dev/null || :
    ++		wait $server_pid 2>/dev/null || :
    ++		exec 7>&-
    ++		exec 8>&-
    ++		rm -f server-ready first-ready slow-pack-server
    ++	" &&
    ++	read port <&7 &&
    ++	url="http://127.0.0.1:$port/pack" &&
    ++	{
    ++		(
    ++			if ! GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-first.trace" \
    ++			GIT_TRACE_CURL_NO_DATA=1 \
    ++			git -C packfileclient-overlap http-fetch --packfile="$packhash" \
    ++				--index-pack-arg=index-pack \
    ++				--index-pack-arg=--stdin --index-pack-arg=--keep \
    ++				"$url" >first.out
    ++			then
    ++				echo failed >"$TRASH_DIRECTORY/first-ready" &&
    ++				exit 1
    ++			fi
    ++		) &
    ++		first_pid=$!
    ++	} &&
    ++	test_when_finished "
    ++		kill $first_pid 2>/dev/null || :
    ++		wait $first_pid 2>/dev/null || :
    ++	" &&
    ++	read ready <&8 &&
    ++	test "$ready" = ready &&
    ++	test_path_is_file "$tmpfile" &&
    ++	test -s "$tmpfile" &&
    ++	{
    ++		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
    ++		GIT_TRACE_CURL_NO_DATA=1 \
    ++		git -C packfileclient-overlap http-fetch --packfile="$packhash" \
    ++			--index-pack-arg=index-pack \
    ++			--index-pack-arg=--stdin --index-pack-arg=--keep \
    ++			"$url" >second.out &
    ++		second_pid=$!
    ++	} &&
    ++	test_when_finished "
    ++		kill $second_pid 2>/dev/null || :
    ++		wait $second_pid 2>/dev/null || :
    ++	" &&
    ++	wait "$server_pid" &&
    ++	wait "$first_pid" &&
    ++	wait "$second_pid" &&
    ++	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
    ++	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
    ++	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
    ++	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
    ++	sort first.out second.out >actual &&
    ++	test_cmp expect actual &&
    ++	test_path_is_missing "$tmpfile" &&
    ++	git -C packfileclient-overlap cat-file -e "$blob"
    ++'
     +
      test_expect_success 'fetch notices corrupt pack' '
      	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
      	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
    -@@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile with corrupt pack' '
    - 	git init packfileclient &&
    - 	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git && ls objects/pack/pack-*.pack) &&
    - 	test_must_fail git -C packfileclient http-fetch --packfile \
    --		"$HTTPD_URL"/dumb/repo_bad1.git/$p
    -+		"$HTTPD_URL"/dumb/repo_bad1.git/$p &&
    -+	find packfileclient/.git/objects/pack -name "tmp_pack_*" -print >tmpfiles &&
    -+	test_must_be_empty tmpfiles
    - '
    - 
    - test_expect_success 'fetch notices corrupt idx' '
2:  e73de423f0 ! 2:  9b41d4ddb3 fetch-pack: accept "pack" output for packfile URIs
    @@ Metadata
      ## Commit message ##
         fetch-pack: accept "pack" output for packfile URIs
     
    -    When "index-pack --keep" creates a .keep file, it reports
    -    "keep<TAB><hash>". If the file already exists, index-pack leaves it
    -    untouched and reports "pack<TAB><hash>" instead.
    +    When index-pack finds an existing keep file it reports pack rather than
    +    keep. Accept either result from http-fetch, and only register a keep
    +    lockfile when this fetch created it.
     
    -    Since dd4b732df7 (upload-pack: send part of packfile response as uri,
    -    2020-06-10), fetch-pack has accepted only the "keep" form for packs
    -    downloaded through packfile URIs. A concurrent fetch can install the
    -    same pack and create its .keep file before another process reaches
    -    index-pack. The latter process then fails even though index-pack
    -    completed successfully.
    -
    -    Accept both successful forms. Add a path to pack_lockfiles only for the
    -    "keep" form, so cleanup removes only a keep file created by the current
    -    process and preserves a pre-existing one.
    -
    -    Add a regression test which pre-creates a keep file and verifies that a
    -    fetch succeeds without changing it.
    +    Read the pack/keep prefix and hash without consuming any following fsck
    +    output, validate the reported pack hash against the advertised hash, and
    +    exercise a packfile URI fetch with a pre-existing keep file.
     
         Signed-off-by: Ted Nyman <tnyman@openai.com>
    -    Signed-off-by: Junio C Hamano <gitster@pobox.com>
     
      ## fetch-pack.c ##
     @@ fetch-pack.c: static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
      	}
      
      	for (i = 0; i < packfile_uris.nr; i++) {
    -+		int created_keep = 0;
    ++		bool created_keep;
      		int j;
      		struct child_process cmd = CHILD_PROCESS_INIT;
     -		char packname[GIT_MAX_HEXSZ + 1];
    -+		char packname[GIT_MAX_HEXSZ + 6];
    -+		const char *packhash;
    -+		const int packname_len = the_hash_algo->hexsz + 6;
    ++		char packhash[GIT_MAX_HEXSZ + 1];
      		const char *uri = packfile_uris.items[i].string +
      			the_hash_algo->hexsz + 1;
      
    @@ fetch-pack.c: static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
     -		if (read_in_full(cmd.out, packname, 5) < 0 ||
     -		    memcmp(packname, "keep\t", 5))
     -			die("fetch-pack: expected keep then TAB at start of http-fetch output");
    --
    ++		if (read_in_full(cmd.out, packhash, 5) != 5 ||
    ++		    (memcmp(packhash, "keep\t", 5) &&
    ++		     memcmp(packhash, "pack\t", 5)))
    ++			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
    ++		created_keep = !memcmp(packhash, "keep\t", 5);
    + 
     -		if (read_in_full(cmd.out, packname,
     -				 the_hash_algo->hexsz + 1) < 0 ||
     -		    packname[the_hash_algo->hexsz] != '\n')
     -			die("fetch-pack: expected hash then LF at end of http-fetch output");
     -
     -		packname[the_hash_algo->hexsz] = '\0';
    -+		if (read_in_full(cmd.out, packname, packname_len) != packname_len ||
    -+		    packname[packname_len - 1] != '\n')
    -+			die("fetch-pack: expected pack or keep, TAB, hash, "
    -+			    "then LF in http-fetch output");
    -+		packname[packname_len - 1] = '\0';
    -+		if (skip_prefix(packname, "keep\t", &packhash))
    -+			created_keep = 1;
    -+		else if (!skip_prefix(packname, "pack\t", &packhash))
    -+			die("fetch-pack: expected pack or keep, TAB, hash, "
    -+			    "then LF in http-fetch output");
    ++		if (read_in_full(cmd.out, packhash,
    ++				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
    ++		    packhash[the_hash_algo->hexsz] != '\n')
    ++			die("fetch-pack: expected hash then LF in http-fetch output");
    ++		packhash[the_hash_algo->hexsz] = '\0';
      
      		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
      
base-commit: f60db8d575adb79761d363e026fb49bddf330c73
-- 
2.55.0.125.g9b41d4ddb3
Ted NymanJul 20, 2026, 22:33 UTC in reply to Ted Nyman on lore

[PATCH v2 1/2] http: avoid concurrent appends to partial packs

Pack requests stage downloads in a predictable partial-pack file so an interrupted transfer can be resumed. Both packfile URI and ordinary dumb HTTP requests use this staging path. Opening it in append mode lets concurrent fetches interleave their writes, corrupting the pack or causing a later fetch to request a range at EOF.

Open the partial pack read-write, seek to its current end, and retain a per-descriptor offset for incoming data. Reopen newly created partial packs without O_CREAT so Windows permits concurrent unlink, and keep the descriptor for index-pack when another downloader removes the staging path. Accept HTTP 416 when a partial pack is already complete.

Exercise resumed transfers, EOF ranges, and overlapping 200 and 206 responses. Clarify the staging-key documentation and correct the stale --index-pack-args spelling in the documentation and error messages; the repeatable --index-pack-arg option is already accepted.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc |  13 +-
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  53 ++++---
 t/t5550-http-fetch-dumb.sh        | 223 ++++++++++++++++++++++++++++++
 6 files changed, 271 insertions(+), 31 deletions(-)
Show changes to 6 files +271 −30

Documentation/git-http-fetch.adoc, http-fetch.c, http-push.c, http-walker.c, http.c, t/t5550-http-fetch-dumb.sh

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 2200f073c4..60ca91cf3a 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -48,13 +48,14 @@ commit-id::
 	line (which is not expected in
 	this case), 'git http-fetch' fetches the packfile directly at the given
 	URL and uses index-pack to generate corresponding .idx and .keep files.
-	The hash is used to determine the name of the temporary file and is
-	arbitrary. The output of index-pack is printed to stdout. Requires
-	--index-pack-args.
+	The hash is used to determine the name of the temporary file. It need
+	not be the pack hash, but it must uniquely identify the pack contents
+	for resumption. The output of index-pack is printed to stdout. Requires
+	one or more --index-pack-arg options.
 
---index-pack-args=<args>::
-	For internal use only. The command to run on the contents of the
-	downloaded pack. Arguments are URL-encoded separated by spaces.
+--index-pack-arg=<arg>::
+	For internal use only. An argument to the command run on the contents
+	of the downloaded pack. This option can be specified multiple times.
 
 --recover::
 	Verify that everything reachable from target is fetched.  Used after
diff --git a/http-fetch.c b/http-fetch.c
index f9b6ecb061..05f68f306a 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			struct url_info url;
 			char *nurl = url_normalize(preq->url, &url);
 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
@@ -155,7 +156,7 @@ int cmd_main(int argc, const char **argv)
 
 	if (packfile) {
 		if (!index_pack_args.nr)
-			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
+			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
 
 		fetch_single_packfile(&packfile_hash, argv[arg],
 				      index_pack_args.v);
@@ -164,7 +165,7 @@ int cmd_main(int argc, const char **argv)
 	}
 
 	if (index_pack_args.nr)
-		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
+		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
 
 	if (commits_on_stdin) {
 		commits = walker_targets_stdin(&commit_id, &write_ref);
diff --git a/http-push.c b/http-push.c
index 3c23cbba27..03dc8102a1 100644
--- a/http-push.c
+++ b/http-push.c
@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)
 
 	} else if (request->state == RUN_FETCH_PACKED) {
 		int fail = 1;
-		if (request->curl_result != CURLE_OK) {
+		if (request->curl_result != CURLE_OK &&
+		    request->http_code != 416) {
 			fprintf(stderr, "Unable to get pack file %s\n%s",
 				request->url, curl_errorstr);
 		} else {
diff --git a/http-walker.c b/http-walker.c
index b58a3b2a92..abafca84d6 100644
--- a/http-walker.c
+++ b/http-walker.c
@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			error("Unable to get pack file %s\n%s", preq->url,
 			      curl_errorstr);
 			goto abort;
diff --git a/http.c b/http.c
index b4e7b8d00b..9b9f4efe28 100644
--- a/http.c
+++ b/http.c
@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	int tmpfile_fd;
 	int ret = 0;
 
+	/* Another downloader may unlink the staging path while we index it. */
+	tmpfile_fd = xdup(fileno(preq->packfile));
 	fclose(preq->packfile);
 	preq->packfile = NULL;
-
-	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
+	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
+		die_errno("unable to seek local file %s for pack",
+			  preq->tmpfile.buf);
 
 	ip.git_cmd = 1;
 	ip.in = tmpfile_fd;
@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	else
 		ip.no_stdout = 1;
 
-	if (run_command(&ip)) {
+	if (run_command(&ip))
 		ret = -1;
-		goto cleanup;
-	}
-
-cleanup:
-	close(tmpfile_fd);
 	unlink(preq->tmpfile.buf);
 	return ret;
 }
@@ -2738,22 +2736,42 @@ struct http_pack_request *new_http_pack_request(
 struct http_pack_request *new_direct_http_pack_request(
 	const unsigned char *packed_git_hash, char *url)
 {
-	off_t prev_posn = 0;
+	off_t prev_posn;
 	struct http_pack_request *preq;
+	int fd;
 
 	CALLOC_ARRAY(preq, 1);
 	strbuf_init(&preq->tmpfile, 0);
-
 	preq->url = url;
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
-	if (!preq->packfile) {
-		error("Unable to open local file %s for pack",
-		      preq->tmpfile.buf);
+	/* Reopen without O_CREAT so MinGW permits another writer to unlink it. */
+	for (;;) {
+		fd = open(preq->tmpfile.buf, O_RDWR);
+		if (fd >= 0 || errno != ENOENT)
+			break;
+		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
+		if (fd >= 0) {
+			close(fd);
+			continue;
+		}
+		if (errno != EEXIST)
+			break;
+	}
+	if (fd < 0) {
+		error_errno("unable to open local file %s for pack",
+			    preq->tmpfile.buf);
+		goto abort;
+	}
+	prev_posn = lseek(fd, 0, SEEK_END);
+	if (prev_posn < 0) {
+		error_errno("unable to seek local file %s for pack",
+			    preq->tmpfile.buf);
+		close(fd);
 		goto abort;
 	}
+	preq->packfile = xfdopen(fd, "w");
 
 	preq->slot = get_active_slot();
 	preq->headers = object_request_headers();
@@ -2762,12 +2780,7 @@ struct http_pack_request *new_direct_http_pack_request(
 	curl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);
 	curl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);
 
-	/*
-	 * If there is data present from a previous transfer attempt,
-	 * resume where it left off
-	 */
-	prev_posn = ftello(preq->packfile);
-	if (prev_posn>0) {
+	if (prev_posn > 0) {
 		if (http_is_verbose)
 			fprintf(stderr,
 				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index b0080bf204..7acae96a96 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -293,6 +293,229 @@ test_expect_success 'http-fetch --packfile' '
 	git -C packfileclient cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile resumes a partial download' '
+	git init packfileclient-resume &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/resume.trace" \
+	git -C packfileclient-resume http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=index-pack --index-pack-arg=--stdin \
+		--index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_grep "Range: bytes=64-" resume.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-resume cat-file -e "$HASH"
+'
+
+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
+	git init packfileclient-concurrent &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	packhash=$(basename "$p" .pack) &&
+	packhash=${packhash#pack-} &&
+	tmpfile="packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	mkfifo first-ready first-continue &&
+	exec 8<>first-ready &&
+	exec 9<>first-continue &&
+	write_script git-wait-index-pack <<-\EOF &&
+	echo ready >"$GIT_TEST_WAIT_READY" &&
+	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
+	exec git index-pack "$@"
+	EOF
+	{
+		(
+			if ! PATH="$TRASH_DIRECTORY:$PATH" \
+			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
+			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
+			GIT_TRACE_CURL="$TRASH_DIRECTORY/first.trace" \
+			git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+				--index-pack-arg=wait-index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		echo continue >&9
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+		exec 8>&-
+		exec 9>&-
+		rm -f first-ready first-continue git-wait-index-pack
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/second.trace" \
+	git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+		--index-pack-arg=index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
+	echo continue >&9 &&
+	wait "$first_pid" &&
+	printf "pack\t%s\n" "$packhash" >expect &&
+	test_cmp expect first.out &&
+	printf "keep\t%s\n" "$packhash" >expect &&
+	test_cmp expect second.out &&
+	test_grep "Range: bytes=64-" first.trace &&
+	test_grep "Range: bytes=[0-9]*-" second.trace &&
+	test_grep "HTTP/[0-9.]* 416" second.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-concurrent cat-file -e "$HASH"
+'
+
+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
+	git init packfileclient-overlap &&
+	blob=$(test-tool genrandom pack-overlap 2m |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			hash-object -w --stdin) &&
+	packhash=$(printf "%s\n" "$blob" |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			pack-objects "$TRASH_DIRECTORY/overlap-pack") &&
+	pack="$TRASH_DIRECTORY/overlap-pack-$packhash.pack" &&
+	tmpfile="packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp" &&
+	mkfifo server-ready first-ready &&
+	exec 7<>server-ready &&
+	exec 8<>first-ready &&
+	write_script slow-pack-server "$PERL_PATH" <<-\EOF &&
+	use strict;
+	use warnings;
+	use IO::Socket::INET;
+
+	my ($packfile, $server_ready, $first_ready) = @ARGV;
+	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
+	my $pack = do { local $/; <$in> };
+	close($in) or die "close $packfile: $!";
+	my $server = IO::Socket::INET->new(LocalAddr => "127.0.0.1",
+		LocalPort => 0, Proto => "tcp", Listen => 2, ReuseAddr => 1)
+		or die "listen: $!";
+
+	sub signal_ready {
+		my ($file, $value) = @_;
+		open(my $out, ">", $file) or die "open $file: $!";
+		print $out "$value\n" or die "write $file: $!";
+		close($out) or die "close $file: $!";
+	}
+
+	sub write_all {
+		my ($out, $data) = @_;
+		my $offset = 0;
+		while ($offset < length($data)) {
+			my $written = syswrite($out, $data,
+				length($data) - $offset, $offset);
+			defined($written) && $written or die "write response: $!";
+			$offset += $written;
+		}
+	}
+
+	sub start_response {
+		my $out = $server->accept() or die "accept: $!";
+		<$out> or die "read request: $!";
+		my $start = 0;
+		while (<$out>) {
+			last if /^\r?\n$/;
+			$start = $1 if /^Range: bytes=(\d+)-/i;
+		}
+		$start < length($pack) or die "invalid range $start";
+		my $length = length($pack) - $start;
+		my $middle = int($length / 2);
+		my $status = $start ? "206 Partial Content" : "200 OK";
+		my $headers = "HTTP/1.1 $status\r\n" .
+			"Content-Length: $length\r\n" .
+			($start ? "Content-Range: bytes $start-" .
+				(length($pack) - 1) . "/" . length($pack) . "\r\n" : "") .
+			"Connection: close\r\n\r\n";
+		write_all($out, $headers);
+		write_all($out, substr($pack, $start, $middle));
+		return ($out, $start + $middle);
+	}
+
+	signal_ready($server_ready, $server->sockport());
+	my ($first, $first_pos) = start_response();
+	signal_ready($first_ready, "ready");
+	my ($second, $second_pos) = start_response();
+	write_all($first, substr($pack, $first_pos));
+	write_all($second, substr($pack, $second_pos));
+	close($first) or die "close first response: $!";
+	close($second) or die "close second response: $!";
+	EOF
+	{
+		(
+			if ! "$TRASH_DIRECTORY/slow-pack-server" "$pack" \
+				"$TRASH_DIRECTORY/server-ready" \
+				"$TRASH_DIRECTORY/first-ready"
+			then
+				echo failed >"$TRASH_DIRECTORY/server-ready" &&
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) >server.log 2>&1 &
+		server_pid=$!
+	} &&
+	test_when_finished "
+		kill $server_pid 2>/dev/null || :
+		wait $server_pid 2>/dev/null || :
+		exec 7>&-
+		exec 8>&-
+		rm -f server-ready first-ready slow-pack-server
+	" &&
+	read port <&7 &&
+	url="http://127.0.0.1:$port/pack" &&
+	{
+		(
+			if ! GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-first.trace" \
+			GIT_TRACE_CURL_NO_DATA=1 \
+			git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+				--index-pack-arg=index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$url" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	test_path_is_file "$tmpfile" &&
+	test -s "$tmpfile" &&
+	{
+		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
+		GIT_TRACE_CURL_NO_DATA=1 \
+		git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+			--index-pack-arg=index-pack \
+			--index-pack-arg=--stdin --index-pack-arg=--keep \
+			"$url" >second.out &
+		second_pid=$!
+	} &&
+	test_when_finished "
+		kill $second_pid 2>/dev/null || :
+		wait $second_pid 2>/dev/null || :
+	" &&
+	wait "$server_pid" &&
+	wait "$first_pid" &&
+	wait "$second_pid" &&
+	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
+	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
+	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
+	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
+	sort first.out second.out >actual &&
+	test_cmp expect actual &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-overlap cat-file -e "$blob"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
-- 
2.55.0.125.g9b41d4ddb3
Ted NymanJul 20, 2026, 22:34 UTC in reply to Ted Nyman on lore

[PATCH v2 2/2] fetch-pack: accept "pack" output for packfile URIs

When index-pack finds an existing keep file it reports pack rather than keep. Accept either result from http-fetch, and only register a keep lockfile when this fetch created it.

Read the pack/keep prefix and hash without consuming any following fsck output, validate the reported pack hash against the advertised hash, and exercise a packfile URI fetch with a pre-existing keep file.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 fetch-pack.c           | 33 ++++++++++++++++++---------------
 t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
 2 files changed, 49 insertions(+), 15 deletions(-)
Show changes to 2 files +49 −15

fetch-pack.c, t/t5702-protocol-v2.sh

diff --git a/fetch-pack.c b/fetch-pack.c
index 120e01f3cf..509b91527b 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 	}
 
 	for (i = 0; i < packfile_uris.nr; i++) {
+		bool created_keep;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 1];
+		char packhash[GIT_MAX_HEXSZ + 1];
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, 5) < 0 ||
-		    memcmp(packname, "keep\t", 5))
-			die("fetch-pack: expected keep then TAB at start of http-fetch output");
+		if (read_in_full(cmd.out, packhash, 5) != 5 ||
+		    (memcmp(packhash, "keep\t", 5) &&
+		     memcmp(packhash, "pack\t", 5)))
+			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
+		created_keep = !memcmp(packhash, "keep\t", 5);
 
-		if (read_in_full(cmd.out, packname,
-				 the_hash_algo->hexsz + 1) < 0 ||
-		    packname[the_hash_algo->hexsz] != '\n')
-			die("fetch-pack: expected hash then LF at end of http-fetch output");
-
-		packname[the_hash_algo->hexsz] = '\0';
+		if (read_in_full(cmd.out, packhash,
+				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
+		    packhash[the_hash_algo->hexsz] != '\n')
+			die("fetch-pack: expected hash then LF in http-fetch output");
+		packhash[the_hash_algo->hexsz] = '\0';
 
 		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
 
@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (finish_command(&cmd))
 			die("fetch-pack: unable to finish http-fetch");
 
-		if (memcmp(packfile_uris.items[i].string, packname,
+		if (memcmp(packfile_uris.items[i].string, packhash,
 			   the_hash_algo->hexsz))
 			die("fetch-pack: pack downloaded from %s does not match expected hash %.*s",
 			    uri, (int) the_hash_algo->hexsz,
 			    packfile_uris.items[i].string);
 
-		string_list_append_nodup(pack_lockfiles,
-					 xstrfmt("%s/pack/pack-%s.keep",
-						 repo_get_object_directory(the_repository),
-						 packname));
+		if (created_keep)
+			string_list_append_nodup(pack_lockfiles,
+						 xstrfmt("%s/pack/pack-%s.keep",
+							 repo_get_object_directory(the_repository),
+							 packhash));
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);
diff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh
index 9f6cf4142d..1861eb7d7c 100755
--- a/t/t5702-protocol-v2.sh
+++ b/t/t5702-protocol-v2.sh
@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '
 		fetch "$HTTPD_URL/smart/http_parent"
 '
 
+test_expect_success 'packfile URI preserves an existing keep file' '
+	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
+	rm -rf "$P" http_child keep.expect &&
+
+	git init "$P" &&
+	git -C "$P" config uploadpack.allowsidebandall true &&
+
+	echo my-blob >"$P/my-blob" &&
+	git -C "$P" add my-blob &&
+	git -C "$P" commit -m x &&
+	configure_exclusion "$P" my-blob >h &&
+
+	git init http_child &&
+	packhash=$(cat packh) &&
+	keep="http_child/.git/objects/pack/pack-$packhash.keep" &&
+	echo pre-existing >"$keep" &&
+	cp "$keep" keep.expect &&
+
+	GIT_TEST_SIDEBAND_ALL=1 \
+	git -C http_child -c protocol.version=2 \
+		-c fetch.uriprotocols=http,https \
+		fetch "$HTTPD_URL/smart/http_parent" &&
+
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.pack" &&
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.idx" &&
+	test_cmp keep.expect "$keep" &&
+	git -C http_child cat-file -e "$(cat h)"
+'
+
 test_expect_success 'fetching with valid packfile URI but invalid hash fails' '
 	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
 	rm -rf "$P" http_child log &&
-- 
2.55.0.125.g9b41d4ddb3
Junio C HamanoJul 21, 2026, 19:56 UTC in reply to Ted Nyman on lore

Re: [PATCH v2 1/2] http: avoid concurrent appends to partial packs

Ted Nyman <tnyman@openai.com> writes:
Show 16 quoted lines
> Pack requests stage downloads in a predictable partial-pack file so an
> interrupted transfer can be resumed. Both packfile URI and ordinary dumb
> HTTP requests use this staging path. Opening it in append mode lets
> concurrent fetches interleave their writes, corrupting the pack or
> causing a later fetch to request a range at EOF.
>
> Open the partial pack read-write, seek to its current end, and retain a
> per-descriptor offset for incoming data. Reopen newly created partial
> packs without O_CREAT so Windows permits concurrent unlink, and keep the
> descriptor for index-pack when another downloader removes the staging
> path. Accept HTTP 416 when a partial pack is already complete.
>
> Exercise resumed transfers, EOF ranges, and overlapping 200 and 206
> responses. Clarify the staging-key documentation and correct the stale
> --index-pack-args spelling in the documentation and error messages; the
> repeatable --index-pack-arg option is already accepted.

Hmph. So the idea is to allow multiple processes to open the same file and, because they all know where their respective chunks of data fit in the final file, have them use pwrite(2) to deposit those pieces at the exact target locations, and this prevents them from stepping on each other's toes?

I cannot exactly explain why but it somehow makes me feel dirty.

It is also surprising that the workaround on MinGW works when one of these multiple processes finishes writing and attempts to finalize the temporary file while others still have open file descriptors to the same file.

Show 7 quoted lines
> -	The hash is used to determine the name of the temporary file and is
> -	arbitrary. The output of index-pack is printed to stdout. Requires
> -	--index-pack-args.
> +	The hash is used to determine the name of the temporary file. It need
> +	not be the pack hash, but it must uniquely identify the pack contents
> +	for resumption. The output of index-pack is printed to stdout. Requires
> +	one or more --index-pack-arg options.
OK.
Show 6 quoted lines
> ---index-pack-args=<args>::
> -	For internal use only. The command to run on the contents of the
> -	downloaded pack. Arguments are URL-encoded separated by spaces.
> +--index-pack-arg=<arg>::
> +	For internal use only. An argument to the command run on the contents
> +	of the downloaded pack. This option can be specified multiple times.

Was the 'internal use only' thing renamed in order to prevent the new code from accidentally working with an older caller?

    ... goes and notices that the code uses singular form throughout ...

Ah, no, this is an unrelated typo fix that remains valid even if the rest of this patch is dropped. Good catch.

It would be easier to review the actual changes if this cleanup were isolated in a preliminary patch. Are there other cleanup changes in this series that fall into the same category?

Show 11 quoted lines
> diff --git a/http-fetch.c b/http-fetch.c
> index f9b6ecb061..05f68f306a 100644
> --- a/http-fetch.c
> +++ b/http-fetch.c
> @@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
>  
>  	if (start_active_slot(preq->slot)) {
>  		run_active_slot(preq->slot);
> -		if (results.curl_result != CURLE_OK) {
> +		if (results.curl_result != CURLE_OK &&
> +		    results.http_code != 416) {

We do not seem to use symbolic constants for these '4xx' codes (or '2xx', for that matter), so I will let that pass. Eventually, we may want to give symbolic constants to them to improve readability, but doing so is certainly outside the scope of this topic.

Show 6 quoted lines
> @@ -155,7 +156,7 @@ int cmd_main(int argc, const char **argv)
>  
>  	if (packfile) {
>  		if (!index_pack_args.nr)
> -			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
> +			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
This and ...
Show 9 quoted lines
> @@ -164,7 +165,7 @@ int cmd_main(int argc, const char **argv)
>  	}
>  
>  	if (index_pack_args.nr)
> -		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
> +		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
>  
>  	if (commits_on_stdin) {
>  		commits = walker_targets_stdin(&commit_id, &write_ref);

... this is the same "index-pack-arg" fix and can be moved to a separate preliminary clean-up patch.

Thanks.
Ted NymanJul 21, 2026, 23:29 UTC in reply to Ted Nyman on lore

[PATCH v3 0/3] packfile URIs: support concurrent downloads

Packfile URI and dumb HTTP downloads stage packs at objects/pack/pack-<hash>.pack.temp so an interrupted transfer can resume. Opening that file in append mode forces every write to its current end. Two Git processes fetching the same pack into one object database can therefore append duplicate data and corrupt the pack.

The first patch separates the unrelated --index-pack-arg documentation and error-message correction requested during review.

The second patch keeps the predictable staging name but removes append mode. Each downloader seeks once to the current end, requests the corresponding Range, and writes using its own descriptor offset. Since the staging key must identify immutable pack contents, overlapping responses write identical bytes at identical offsets. There is no need for pwrite(2) or cross-process coordination, and resumption continues to work for both packfile URI and ordinary dumb HTTP downloads.

A downloader can also find that the partial pack has completed and request a range starting at EOF. Servers may respond with HTTP 416 in that case. Treat the response as a completed download and let index-pack validate the pack.

On MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing staging file exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the path. Keep the open descriptor for index-pack; it installs its own pack, so the shared staging file is only unlinked, never renamed.

The third patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process. Read only the prefix and hash so any following fsck output remains available to fetch-pack.

The tests cover resumption, a completed partial returning 416, overlapping 200 and 206 responses, unlinking the staging path while index-pack holds its descriptor, and a pre-existing .keep file. The unlink test does not require FIFOs, so it can exercise MinGW's sharing behavior even though the concurrent-download tests are skipped there.

Changes since v2:
  * Split the --index-pack-arg documentation and error-message cleanup
    into a preliminary patch, as requested by Junio.
  * Clarify why per-descriptor offsets keep overlapping writes safe and
    why MinGW permits the shared staging path to be unlinked.
  * Add a non-FIFO unlink-while-indexing regression test that can run on
    MinGW.
  * Rebase onto the current master.
The v2 discussion is at:
  https://lore.kernel.org/git/cover.1784582665.git.tnyman@openai.com/
Ted Nyman (3):
  http-fetch: correct --index-pack-arg documentation
  http: avoid concurrent appends to partial packs
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  13 +-
 fetch-pack.c                      |  33 ++--
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 244 ++++++++++++++++++++++++++++++
 t/t5702-protocol-v2.sh            |  31 ++++
 8 files changed, 344 insertions(+), 46 deletions(-)
Range-diff against v2:
-:  ---------- > 1:  a6a40b8046 http-fetch: correct --index-pack-arg documentation
1:  160a9b9fd0 ! 2:  6c91054afc http: avoid concurrent appends to partial packs
    @@ Commit message
     
         Pack requests stage downloads in a predictable partial-pack file so an
         interrupted transfer can be resumed. Both packfile URI and ordinary dumb
    -    HTTP requests use this staging path. Opening it in append mode lets
    -    concurrent fetches interleave their writes, corrupting the pack or
    -    causing a later fetch to request a range at EOF.
    +    HTTP requests use this staging path. Opening it in append mode forces
    +    each write to the current end of the file, so concurrent responses can
    +    append duplicate data and corrupt the pack.
     
    -    Open the partial pack read-write, seek to its current end, and retain a
    -    per-descriptor offset for incoming data. Reopen newly created partial
    -    packs without O_CREAT so Windows permits concurrent unlink, and keep the
    -    descriptor for index-pack when another downloader removes the staging
    -    path. Accept HTTP 416 when a partial pack is already complete.
    +    Open the partial pack read-write without O_APPEND and seek once to its
    +    current end. Each downloader then retains the offset matching the Range
    +    it requested. Because the staging key must uniquely identify immutable
    +    pack contents, overlapping responses write the same bytes at the same
    +    offsets instead of extending the file with duplicate data.
     
    -    Exercise resumed transfers, EOF ranges, and overlapping 200 and 206
    -    responses. Clarify the staging-key documentation and correct the stale
    -    --index-pack-args spelling in the documentation and error messages; the
    -    repeatable --index-pack-arg option is already accepted.
    +    MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
    +    existing file. Create a missing partial pack exclusively, close it, and
    +    reopen it without O_CREAT so every retained descriptor permits another
    +    downloader to unlink the staging path. Duplicate that descriptor for
    +    index-pack instead of reopening the path after closing the stream;
    +    index-pack installs its own pack and the shared staging file is only
    +    unlinked, never renamed. Accept HTTP 416 when a partial pack is already
    +    complete and let index-pack validate its contents.
    +
    +    Exercise resumed transfers, EOF ranges, overlapping 200 and 206
    +    responses, and unlinking the staging path while index-pack still holds
    +    its descriptor. Clarify the staging-key documentation.
     
         Signed-off-by: Ted Nyman <tnyman@openai.com>
     
    @@ Documentation/git-http-fetch.adoc: commit-id::
      	URL and uses index-pack to generate corresponding .idx and .keep files.
     -	The hash is used to determine the name of the temporary file and is
     -	arbitrary. The output of index-pack is printed to stdout. Requires
    --	--index-pack-args.
     +	The hash is used to determine the name of the temporary file. It need
     +	not be the pack hash, but it must uniquely identify the pack contents
     +	for resumption. The output of index-pack is printed to stdout. Requires
    -+	one or more --index-pack-arg options.
    - 
    ----index-pack-args=<args>::
    --	For internal use only. The command to run on the contents of the
    --	downloaded pack. Arguments are URL-encoded separated by spaces.
    -+--index-pack-arg=<arg>::
    -+	For internal use only. An argument to the command run on the contents
    -+	of the downloaded pack. This option can be specified multiple times.
    + 	one or more --index-pack-arg options.
      
    - --recover::
    - 	Verify that everything reachable from target is fetched.  Used after
    + --index-pack-arg=<arg>::
     
      ## http-fetch.c ##
     @@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,
    @@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,
      			struct url_info url;
      			char *nurl = url_normalize(preq->url, &url);
      			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
    -@@ http-fetch.c: int cmd_main(int argc, const char **argv)
    - 
    - 	if (packfile) {
    - 		if (!index_pack_args.nr)
    --			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
    -+			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
    - 
    - 		fetch_single_packfile(&packfile_hash, argv[arg],
    - 				      index_pack_args.v);
    -@@ http-fetch.c: int cmd_main(int argc, const char **argv)
    - 	}
    - 
    - 	if (index_pack_args.nr)
    --		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
    -+		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
    - 
    - 	if (commits_on_stdin) {
    - 		commits = walker_targets_stdin(&commit_id, &write_ref);
     
      ## http-push.c ##
     @@ http-push.c: static void finish_request(struct transfer_request *request)
    @@ http.c: struct http_pack_request *new_http_pack_request(
     -	if (!preq->packfile) {
     -		error("Unable to open local file %s for pack",
     -		      preq->tmpfile.buf);
    -+	/* Reopen without O_CREAT so MinGW permits another writer to unlink it. */
    ++	/*
    ++	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
    ++	 * existing file; reopen a newly created file so others may unlink it.
    ++	 */
     +	for (;;) {
     +		fd = open(preq->tmpfile.buf, O_RDWR);
     +		if (fd >= 0 || errno != ENOENT)
    @@ http.c: struct http_pack_request *new_http_pack_request(
     +	if (fd < 0) {
     +		error_errno("unable to open local file %s for pack",
     +			    preq->tmpfile.buf);
    -+		goto abort;
    -+	}
    + 		goto abort;
    + 	}
     +	prev_posn = lseek(fd, 0, SEEK_END);
     +	if (prev_posn < 0) {
     +		error_errno("unable to seek local file %s for pack",
     +			    preq->tmpfile.buf);
     +		close(fd);
    - 		goto abort;
    - 	}
    ++		goto abort;
    ++	}
     +	preq->packfile = xfdopen(fd, "w");
      
      	preq->slot = get_active_slot();
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	git -C packfileclient-resume cat-file -e "$HASH"
     +'
     +
    ++test_expect_success 'http-fetch --packfile permits unlink while indexing' '
    ++	git init packfileclient-unlink &&
    ++	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
    ++		ls objects/pack/pack-*.pack) &&
    ++	tmpfile="packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
    ++	write_script git-unlink-index-pack <<-\EOF &&
    ++	test -f "$GIT_TEST_PACK_TEMP" || exit 1
    ++	rm "$GIT_TEST_PACK_TEMP" || exit 1
    ++	exec git index-pack "$@"
    ++	EOF
    ++	test_when_finished "rm -f git-unlink-index-pack" &&
    ++	PATH="$TRASH_DIRECTORY:$PATH" \
    ++	GIT_TEST_PACK_TEMP="$TRASH_DIRECTORY/$tmpfile" \
    ++	git -C packfileclient-unlink http-fetch --packfile="$ARBITRARY" \
    ++		--index-pack-arg=unlink-index-pack \
    ++		--index-pack-arg=--stdin --index-pack-arg=--keep \
    ++		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
    ++	test_path_is_missing "$tmpfile" &&
    ++	git -C packfileclient-unlink cat-file -e "$HASH"
    ++'
    ++
     +test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
     +	git init packfileclient-concurrent &&
     +	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
2:  9b41d4ddb3 = 3:  1ee5d7e027 fetch-pack: accept "pack" output for packfile URIs
base-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 21, 2026, 23:29 UTC in reply to Ted Nyman on lore

[PATCH v3 1/3] http-fetch: correct --index-pack-arg documentation

The --packfile mode accepts one --index-pack-arg=<arg> option per argument passed to index-pack, but its documentation and option dependency errors still refer to the plural --index-pack-args form.

Correct the spelling and describe the repeatable per-argument form.
Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc | 8 ++++----
 http-fetch.c                      | 4 ++--
 2 files changed, 6 insertions(+), 6 deletions(-)
Show changes to 2 files +6 −5

Documentation/git-http-fetch.adoc, http-fetch.c

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 2200f073c4..09b5d675ee 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -50,11 +50,11 @@ commit-id::
 	URL and uses index-pack to generate corresponding .idx and .keep files.
 	The hash is used to determine the name of the temporary file and is
 	arbitrary. The output of index-pack is printed to stdout. Requires
-	--index-pack-args.
+	one or more --index-pack-arg options.
 
---index-pack-args=<args>::
-	For internal use only. The command to run on the contents of the
-	downloaded pack. Arguments are URL-encoded separated by spaces.
+--index-pack-arg=<arg>::
+	For internal use only. An argument to the command run on the contents
+	of the downloaded pack. This option can be specified multiple times.
 
 --recover::
 	Verify that everything reachable from target is fetched.  Used after
diff --git a/http-fetch.c b/http-fetch.c
index f9b6ecb061..601a77c3c1 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)
 
 	if (packfile) {
 		if (!index_pack_args.nr)
-			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
+			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
 
 		fetch_single_packfile(&packfile_hash, argv[arg],
 				      index_pack_args.v);
@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)
 	}
 
 	if (index_pack_args.nr)
-		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
+		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
 
 	if (commits_on_stdin) {
 		commits = walker_targets_stdin(&commit_id, &write_ref);
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 21, 2026, 23:29 UTC in reply to Ted Nyman on lore

[PATCH v3 2/3] http: avoid concurrent appends to partial packs

Pack requests stage downloads in a predictable partial-pack file so an interrupted transfer can be resumed. Both packfile URI and ordinary dumb HTTP requests use this staging path. Opening it in append mode forces each write to the current end of the file, so concurrent responses can append duplicate data and corrupt the pack.

Open the partial pack read-write without O_APPEND and seek once to its current end. Each downloader then retains the offset matching the Range it requested. Because the staging key must uniquely identify immutable pack contents, overlapping responses write the same bytes at the same offsets instead of extending the file with duplicate data.

MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing partial pack exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the staging path. Duplicate that descriptor for index-pack instead of reopening the path after closing the stream; index-pack installs its own pack and the shared staging file is only unlinked, never renamed. Accept HTTP 416 when a partial pack is already complete and let index-pack validate its contents.

Exercise resumed transfers, EOF ranges, overlapping 200 and 206 responses, and unlinking the staging path while index-pack still holds its descriptor. Clarify the staging-key documentation.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc |   5 +-
 http-fetch.c                      |   3 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 244 ++++++++++++++++++++++++++++++
 6 files changed, 289 insertions(+), 25 deletions(-)
Show changes to 6 files +289 −25

Documentation/git-http-fetch.adoc, http-fetch.c, http-push.c, http-walker.c, http.c, t/t5550-http-fetch-dumb.sh

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 09b5d675ee..60ca91cf3a 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -48,8 +48,9 @@ commit-id::
 	line (which is not expected in
 	this case), 'git http-fetch' fetches the packfile directly at the given
 	URL and uses index-pack to generate corresponding .idx and .keep files.
-	The hash is used to determine the name of the temporary file and is
-	arbitrary. The output of index-pack is printed to stdout. Requires
+	The hash is used to determine the name of the temporary file. It need
+	not be the pack hash, but it must uniquely identify the pack contents
+	for resumption. The output of index-pack is printed to stdout. Requires
 	one or more --index-pack-arg options.
 
 --index-pack-arg=<arg>::
diff --git a/http-fetch.c b/http-fetch.c
index 601a77c3c1..05f68f306a 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			struct url_info url;
 			char *nurl = url_normalize(preq->url, &url);
 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
diff --git a/http-push.c b/http-push.c
index 60f6f8f054..ef8abe3908 100644
--- a/http-push.c
+++ b/http-push.c
@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)
 
 	} else if (request->state == RUN_FETCH_PACKED) {
 		int fail = 1;
-		if (request->curl_result != CURLE_OK) {
+		if (request->curl_result != CURLE_OK &&
+		    request->http_code != 416) {
 			fprintf(stderr, "Unable to get pack file %s\n%s",
 				request->url, curl_errorstr);
 		} else {
diff --git a/http-walker.c b/http-walker.c
index b58a3b2a92..abafca84d6 100644
--- a/http-walker.c
+++ b/http-walker.c
@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			error("Unable to get pack file %s\n%s", preq->url,
 			      curl_errorstr);
 			goto abort;
diff --git a/http.c b/http.c
index caccf2108e..a0d399b274 100644
--- a/http.c
+++ b/http.c
@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	int tmpfile_fd;
 	int ret = 0;
 
+	/* Another downloader may unlink the staging path while we index it. */
+	tmpfile_fd = xdup(fileno(preq->packfile));
 	fclose(preq->packfile);
 	preq->packfile = NULL;
-
-	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
+	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
+		die_errno("unable to seek local file %s for pack",
+			  preq->tmpfile.buf);
 
 	ip.git_cmd = 1;
 	ip.in = tmpfile_fd;
@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	else
 		ip.no_stdout = 1;
 
-	if (run_command(&ip)) {
+	if (run_command(&ip))
 		ret = -1;
-		goto cleanup;
-	}
-
-cleanup:
-	close(tmpfile_fd);
 	unlink(preq->tmpfile.buf);
 	return ret;
 }
@@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(
 struct http_pack_request *new_direct_http_pack_request(
 	const unsigned char *packed_git_hash, char *url)
 {
-	off_t prev_posn = 0;
+	off_t prev_posn;
 	struct http_pack_request *preq;
+	int fd;
 
 	CALLOC_ARRAY(preq, 1);
 	strbuf_init(&preq->tmpfile, 0);
-
 	preq->url = url;
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
-	if (!preq->packfile) {
-		error("Unable to open local file %s for pack",
-		      preq->tmpfile.buf);
+	/*
+	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
+	 * existing file; reopen a newly created file so others may unlink it.
+	 */
+	for (;;) {
+		fd = open(preq->tmpfile.buf, O_RDWR);
+		if (fd >= 0 || errno != ENOENT)
+			break;
+		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
+		if (fd >= 0) {
+			close(fd);
+			continue;
+		}
+		if (errno != EEXIST)
+			break;
+	}
+	if (fd < 0) {
+		error_errno("unable to open local file %s for pack",
+			    preq->tmpfile.buf);
 		goto abort;
 	}
+	prev_posn = lseek(fd, 0, SEEK_END);
+	if (prev_posn < 0) {
+		error_errno("unable to seek local file %s for pack",
+			    preq->tmpfile.buf);
+		close(fd);
+		goto abort;
+	}
+	preq->packfile = xfdopen(fd, "w");
 
 	preq->slot = get_active_slot();
 	preq->headers = object_request_headers();
@@ -2762,12 +2783,7 @@ struct http_pack_request *new_direct_http_pack_request(
 	curl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);
 	curl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);
 
-	/*
-	 * If there is data present from a previous transfer attempt,
-	 * resume where it left off
-	 */
-	prev_posn = ftello(preq->packfile);
-	if (prev_posn>0) {
+	if (prev_posn > 0) {
 		if (http_is_verbose)
 			fprintf(stderr,
 				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index f00eeae48f..65b42c4719 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -293,6 +293,250 @@ test_expect_success 'http-fetch --packfile' '
 	git -C packfileclient cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile resumes a partial download' '
+	git init packfileclient-resume &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/resume.trace" \
+	git -C packfileclient-resume http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=index-pack --index-pack-arg=--stdin \
+		--index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_grep "Range: bytes=64-" resume.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-resume cat-file -e "$HASH"
+'
+
+test_expect_success 'http-fetch --packfile permits unlink while indexing' '
+	git init packfileclient-unlink &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	write_script git-unlink-index-pack <<-\EOF &&
+	test -f "$GIT_TEST_PACK_TEMP" || exit 1
+	rm "$GIT_TEST_PACK_TEMP" || exit 1
+	exec git index-pack "$@"
+	EOF
+	test_when_finished "rm -f git-unlink-index-pack" &&
+	PATH="$TRASH_DIRECTORY:$PATH" \
+	GIT_TEST_PACK_TEMP="$TRASH_DIRECTORY/$tmpfile" \
+	git -C packfileclient-unlink http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=unlink-index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-unlink cat-file -e "$HASH"
+'
+
+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
+	git init packfileclient-concurrent &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	packhash=$(basename "$p" .pack) &&
+	packhash=${packhash#pack-} &&
+	tmpfile="packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	mkfifo first-ready first-continue &&
+	exec 8<>first-ready &&
+	exec 9<>first-continue &&
+	write_script git-wait-index-pack <<-\EOF &&
+	echo ready >"$GIT_TEST_WAIT_READY" &&
+	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
+	exec git index-pack "$@"
+	EOF
+	{
+		(
+			if ! PATH="$TRASH_DIRECTORY:$PATH" \
+			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
+			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
+			GIT_TRACE_CURL="$TRASH_DIRECTORY/first.trace" \
+			git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+				--index-pack-arg=wait-index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		echo continue >&9
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+		exec 8>&-
+		exec 9>&-
+		rm -f first-ready first-continue git-wait-index-pack
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/second.trace" \
+	git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+		--index-pack-arg=index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
+	echo continue >&9 &&
+	wait "$first_pid" &&
+	printf "pack\t%s\n" "$packhash" >expect &&
+	test_cmp expect first.out &&
+	printf "keep\t%s\n" "$packhash" >expect &&
+	test_cmp expect second.out &&
+	test_grep "Range: bytes=64-" first.trace &&
+	test_grep "Range: bytes=[0-9]*-" second.trace &&
+	test_grep "HTTP/[0-9.]* 416" second.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-concurrent cat-file -e "$HASH"
+'
+
+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
+	git init packfileclient-overlap &&
+	blob=$(test-tool genrandom pack-overlap 2m |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			hash-object -w --stdin) &&
+	packhash=$(printf "%s\n" "$blob" |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			pack-objects "$TRASH_DIRECTORY/overlap-pack") &&
+	pack="$TRASH_DIRECTORY/overlap-pack-$packhash.pack" &&
+	tmpfile="packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp" &&
+	mkfifo server-ready first-ready &&
+	exec 7<>server-ready &&
+	exec 8<>first-ready &&
+	write_script slow-pack-server "$PERL_PATH" <<-\EOF &&
+	use strict;
+	use warnings;
+	use IO::Socket::INET;
+
+	my ($packfile, $server_ready, $first_ready) = @ARGV;
+	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
+	my $pack = do { local $/; <$in> };
+	close($in) or die "close $packfile: $!";
+	my $server = IO::Socket::INET->new(LocalAddr => "127.0.0.1",
+		LocalPort => 0, Proto => "tcp", Listen => 2, ReuseAddr => 1)
+		or die "listen: $!";
+
+	sub signal_ready {
+		my ($file, $value) = @_;
+		open(my $out, ">", $file) or die "open $file: $!";
+		print $out "$value\n" or die "write $file: $!";
+		close($out) or die "close $file: $!";
+	}
+
+	sub write_all {
+		my ($out, $data) = @_;
+		my $offset = 0;
+		while ($offset < length($data)) {
+			my $written = syswrite($out, $data,
+				length($data) - $offset, $offset);
+			defined($written) && $written or die "write response: $!";
+			$offset += $written;
+		}
+	}
+
+	sub start_response {
+		my $out = $server->accept() or die "accept: $!";
+		<$out> or die "read request: $!";
+		my $start = 0;
+		while (<$out>) {
+			last if /^\r?\n$/;
+			$start = $1 if /^Range: bytes=(\d+)-/i;
+		}
+		$start < length($pack) or die "invalid range $start";
+		my $length = length($pack) - $start;
+		my $middle = int($length / 2);
+		my $status = $start ? "206 Partial Content" : "200 OK";
+		my $headers = "HTTP/1.1 $status\r\n" .
+			"Content-Length: $length\r\n" .
+			($start ? "Content-Range: bytes $start-" .
+				(length($pack) - 1) . "/" . length($pack) . "\r\n" : "") .
+			"Connection: close\r\n\r\n";
+		write_all($out, $headers);
+		write_all($out, substr($pack, $start, $middle));
+		return ($out, $start + $middle);
+	}
+
+	signal_ready($server_ready, $server->sockport());
+	my ($first, $first_pos) = start_response();
+	signal_ready($first_ready, "ready");
+	my ($second, $second_pos) = start_response();
+	write_all($first, substr($pack, $first_pos));
+	write_all($second, substr($pack, $second_pos));
+	close($first) or die "close first response: $!";
+	close($second) or die "close second response: $!";
+	EOF
+	{
+		(
+			if ! "$TRASH_DIRECTORY/slow-pack-server" "$pack" \
+				"$TRASH_DIRECTORY/server-ready" \
+				"$TRASH_DIRECTORY/first-ready"
+			then
+				echo failed >"$TRASH_DIRECTORY/server-ready" &&
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) >server.log 2>&1 &
+		server_pid=$!
+	} &&
+	test_when_finished "
+		kill $server_pid 2>/dev/null || :
+		wait $server_pid 2>/dev/null || :
+		exec 7>&-
+		exec 8>&-
+		rm -f server-ready first-ready slow-pack-server
+	" &&
+	read port <&7 &&
+	url="http://127.0.0.1:$port/pack" &&
+	{
+		(
+			if ! GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-first.trace" \
+			GIT_TRACE_CURL_NO_DATA=1 \
+			git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+				--index-pack-arg=index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$url" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	test_path_is_file "$tmpfile" &&
+	test -s "$tmpfile" &&
+	{
+		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
+		GIT_TRACE_CURL_NO_DATA=1 \
+		git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+			--index-pack-arg=index-pack \
+			--index-pack-arg=--stdin --index-pack-arg=--keep \
+			"$url" >second.out &
+		second_pid=$!
+	} &&
+	test_when_finished "
+		kill $second_pid 2>/dev/null || :
+		wait $second_pid 2>/dev/null || :
+	" &&
+	wait "$server_pid" &&
+	wait "$first_pid" &&
+	wait "$second_pid" &&
+	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
+	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
+	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
+	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
+	sort first.out second.out >actual &&
+	test_cmp expect actual &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-overlap cat-file -e "$blob"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 21, 2026, 23:29 UTC in reply to Ted Nyman on lore

[PATCH v3 3/3] fetch-pack: accept "pack" output for packfile URIs

When index-pack finds an existing keep file it reports pack rather than keep. Accept either result from http-fetch, and only register a keep lockfile when this fetch created it.

Read the pack/keep prefix and hash without consuming any following fsck output, validate the reported pack hash against the advertised hash, and exercise a packfile URI fetch with a pre-existing keep file.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 fetch-pack.c           | 33 ++++++++++++++++++---------------
 t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
 2 files changed, 49 insertions(+), 15 deletions(-)
Show changes to 2 files +49 −15

fetch-pack.c, t/t5702-protocol-v2.sh

diff --git a/fetch-pack.c b/fetch-pack.c
index 29c41132ee..e9f24fbd63 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 	}
 
 	for (i = 0; i < packfile_uris.nr; i++) {
+		bool created_keep;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 1];
+		char packhash[GIT_MAX_HEXSZ + 1];
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, 5) < 0 ||
-		    memcmp(packname, "keep\t", 5))
-			die("fetch-pack: expected keep then TAB at start of http-fetch output");
+		if (read_in_full(cmd.out, packhash, 5) != 5 ||
+		    (memcmp(packhash, "keep\t", 5) &&
+		     memcmp(packhash, "pack\t", 5)))
+			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
+		created_keep = !memcmp(packhash, "keep\t", 5);
 
-		if (read_in_full(cmd.out, packname,
-				 the_hash_algo->hexsz + 1) < 0 ||
-		    packname[the_hash_algo->hexsz] != '\n')
-			die("fetch-pack: expected hash then LF at end of http-fetch output");
-
-		packname[the_hash_algo->hexsz] = '\0';
+		if (read_in_full(cmd.out, packhash,
+				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
+		    packhash[the_hash_algo->hexsz] != '\n')
+			die("fetch-pack: expected hash then LF in http-fetch output");
+		packhash[the_hash_algo->hexsz] = '\0';
 
 		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
 
@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (finish_command(&cmd))
 			die("fetch-pack: unable to finish http-fetch");
 
-		if (memcmp(packfile_uris.items[i].string, packname,
+		if (memcmp(packfile_uris.items[i].string, packhash,
 			   the_hash_algo->hexsz))
 			die("fetch-pack: pack downloaded from %s does not match expected hash %.*s",
 			    uri, (int) the_hash_algo->hexsz,
 			    packfile_uris.items[i].string);
 
-		string_list_append_nodup(pack_lockfiles,
-					 xstrfmt("%s/pack/pack-%s.keep",
-						 repo_get_object_directory(the_repository),
-						 packname));
+		if (created_keep)
+			string_list_append_nodup(pack_lockfiles,
+						 xstrfmt("%s/pack/pack-%s.keep",
+							 repo_get_object_directory(the_repository),
+							 packhash));
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);
diff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh
index 74a2b7730b..0f05286de8 100755
--- a/t/t5702-protocol-v2.sh
+++ b/t/t5702-protocol-v2.sh
@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '
 		fetch "$HTTPD_URL/smart/http_parent"
 '
 
+test_expect_success 'packfile URI preserves an existing keep file' '
+	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
+	rm -rf "$P" http_child keep.expect &&
+
+	git init "$P" &&
+	git -C "$P" config uploadpack.allowsidebandall true &&
+
+	echo my-blob >"$P/my-blob" &&
+	git -C "$P" add my-blob &&
+	git -C "$P" commit -m x &&
+	configure_exclusion "$P" my-blob >h &&
+
+	git init http_child &&
+	packhash=$(cat packh) &&
+	keep="http_child/.git/objects/pack/pack-$packhash.keep" &&
+	echo pre-existing >"$keep" &&
+	cp "$keep" keep.expect &&
+
+	GIT_TEST_SIDEBAND_ALL=1 \
+	git -C http_child -c protocol.version=2 \
+		-c fetch.uriprotocols=http,https \
+		fetch "$HTTPD_URL/smart/http_parent" &&
+
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.pack" &&
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.idx" &&
+	test_cmp keep.expect "$keep" &&
+	git -C http_child cat-file -e "$(cat h)"
+'
+
 test_expect_success 'fetching with valid packfile URI but invalid hash fails' '
 	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
 	rm -rf "$P" http_child log &&
-- 
2.55.0.openai.131.g83a728de1eb6
Junio C HamanoJul 24, 2026, 04:43 UTC in reply to Ted Nyman on lore

Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads

Ted Nyman <tnyman@openai.com> writes:
Show 21 quoted lines
> Packfile URI and dumb HTTP downloads stage packs at
> objects/pack/pack-<hash>.pack.temp so an interrupted transfer can
> resume. Opening that file in append mode forces every write to its
> current end. Two Git processes fetching the same pack into one object
> database can therefore append duplicate data and corrupt the pack.
> ...
> The tests cover resumption, a completed partial returning 416,
> overlapping 200 and 206 responses, unlinking the staging path while
> index-pack holds its descriptor, and a pre-existing .keep file. The
> unlink test does not require FIFOs, so it can exercise MinGW's sharing
> behavior even though the concurrent-download tests are skipped there.
>
> Changes since v2:
>
>   * Split the --index-pack-arg documentation and error-message cleanup
>     into a preliminary patch, as requested by Junio.
>   * Clarify why per-descriptor offsets keep overlapping writes safe and
>     why MinGW permits the shared staging path to be unlinked.
>   * Add a non-FIFO unlink-while-indexing regression test that can run on
>     MinGW.
>   * Rebase onto the current master.

When merged into 'seen', this topic seems to cause t5550 to hang fairly consistently. It is not surprising, considering that the topic adds roughly 240 lines to the test script in question. It is entirely possible that we are seeing an existing breakage from another topic in 'seen' that is exposed by the additional tests.

The CI run
  https://github.com/git/git/actions/runs/30045343889

is today's seen (excluding this topic) at 728e180b7b; it has breakages in leak checking jobs from other topics, but does not see t5550 hanging.

The CI run
  https://github.com/git/git/actions/runs/30048327878

is seen at 05d0dd408c that merges this topic on top of 728e180b7b above. It breaks the same leak checks, but in addition makes t5550 hang.

Can you help figure out what is going on?
Thanks.
PS. Recent CI runs on 'seen' started to spend so much time on static
    analysis (aka coccinelle) jobs, even though I do not think we
    acquired any new rules recently.  We probably need to figure out
    what is going on there, too.  There is something wrong for these
    CI runs that usually take ~40 minutes to spin for more than 4
    hours.
Ted NymanJul 24, 2026, 08:14 UTC in reply to Ted Nyman on lore

[PATCH v4 0/3] packfile URIs: support concurrent downloads

Packfile URI and dumb HTTP downloads stage packs at objects/pack/pack-<hash>.pack.temp so an interrupted transfer can resume. Opening that file in append mode forces every write to its current end. Two Git processes fetching the same pack into one object database can therefore append duplicate data and corrupt the pack.

The first patch separates the unrelated --index-pack-arg documentation and error-message correction requested during review.

The second patch keeps the predictable staging name but removes append mode. Each downloader seeks once to the current end, requests the corresponding Range, and writes using its own descriptor offset. Since the staging key must identify immutable pack contents, overlapping responses write identical bytes at identical offsets. There is no need for pwrite(2) or cross-process coordination, and resumption continues to work for both packfile URI and ordinary dumb HTTP downloads.

A downloader can also find that the partial pack has completed and request a range starting at EOF. Servers may respond with HTTP 416 in that case. Treat the response as a completed download and let index-pack validate the pack.

On MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing staging file exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the path. Keep the open descriptor for index-pack; it installs its own pack, so the shared staging file is only unlinked, never renamed.

The third patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process. Read only the prefix and hash so any following fsck output remains available to fetch-pack.

The tests cover resumption, a completed partial returning 416, overlapping 200 and 206 responses, unlinking the staging path while index-pack holds its descriptor, and a pre-existing .keep file. The unlink test does not require FIFOs, so it can exercise MinGW's sharing behavior even though the concurrent-download tests are skipped there.

Changes since v3:
  * Match HTTP 416 in trace output from both older and current libcurl.
  * Add a timeout to the overlapping-download test server, notify FIFO
    waiters on server failures, and track the actual server process for
    cleanup.
  * Wait for the second downloader first so an early failure cannot
    leave the test server waiting for a request that will never arrive.
  * No production code changes.

These changes avoid false failures with older libcurl and prevent a failed downloader from leaving the test server running indefinitely.

The v3 discussion is at:
  https://lore.kernel.org/git/cover.1784676106.git.tnyman@openai.com/
Ted Nyman (3):
  http-fetch: correct --index-pack-arg documentation
  http: avoid concurrent appends to partial packs
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  13 +-
 fetch-pack.c                      |  33 ++--
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 250 ++++++++++++++++++++++++++++++
 t/t5702-protocol-v2.sh            |  31 ++++
 8 files changed, 350 insertions(+), 46 deletions(-)
Range-diff against v3:
1:  a6a40b8046 = 1:  a6a40b8046 http-fetch: correct --index-pack-arg documentation
2:  6c91054afc ! 2:  144c98cdfa http: avoid concurrent appends to partial packs
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	test_cmp expect second.out &&
     +	test_grep "Range: bytes=64-" first.trace &&
     +	test_grep "Range: bytes=[0-9]*-" second.trace &&
    -+	test_grep "HTTP/[0-9.]* 416" second.trace &&
    ++	test_grep "416 Requested Range Not Satisfiable" second.trace &&
     +	test_path_is_missing "$tmpfile" &&
     +	git -C packfileclient-concurrent cat-file -e "$HASH"
     +'
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	use IO::Socket::INET;
     +
     +	my ($packfile, $server_ready, $first_ready) = @ARGV;
    ++	my $completed = 0;
    ++	END {
    ++		if (!$completed) {
    ++			signal_ready($server_ready, "failed");
    ++			signal_ready($first_ready, "failed");
    ++		}
    ++	}
    ++
    ++	$SIG{ALRM} = sub { die "timed out serving concurrent pack requests\n" };
    ++	alarm 60;
    ++
     +	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
     +	my $pack = do { local $/; <$in> };
     +	close($in) or die "close $packfile: $!";
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	write_all($second, substr($pack, $second_pos));
     +	close($first) or die "close first response: $!";
     +	close($second) or die "close second response: $!";
    ++	$completed = 1;
    ++	alarm 0;
     +	EOF
     +	{
    -+		(
    -+			if ! "$TRASH_DIRECTORY/slow-pack-server" "$pack" \
    -+				"$TRASH_DIRECTORY/server-ready" \
    -+				"$TRASH_DIRECTORY/first-ready"
    -+			then
    -+				echo failed >"$TRASH_DIRECTORY/server-ready" &&
    -+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
    -+				exit 1
    -+			fi
    -+		) >server.log 2>&1 &
    ++		"$TRASH_DIRECTORY/slow-pack-server" "$pack" \
    ++			"$TRASH_DIRECTORY/server-ready" \
    ++			"$TRASH_DIRECTORY/first-ready" >server.log 2>&1 &
     +		server_pid=$!
     +	} &&
     +	test_when_finished "
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +		kill $second_pid 2>/dev/null || :
     +		wait $second_pid 2>/dev/null || :
     +	" &&
    -+	wait "$server_pid" &&
    -+	wait "$first_pid" &&
     +	wait "$second_pid" &&
    ++	wait "$first_pid" &&
    ++	wait "$server_pid" &&
     +	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
     +	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
     +	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
3:  1ee5d7e027 = 3:  d9063deb60 fetch-pack: accept "pack" output for packfile URIs
base-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 24, 2026, 08:14 UTC in reply to Ted Nyman on lore

[PATCH v4 1/3] http-fetch: correct --index-pack-arg documentation

The --packfile mode accepts one --index-pack-arg=<arg> option per argument passed to index-pack, but its documentation and option dependency errors still refer to the plural --index-pack-args form.

Correct the spelling and describe the repeatable per-argument form.
Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc | 8 ++++----
 http-fetch.c                      | 4 ++--
 2 files changed, 6 insertions(+), 6 deletions(-)
Show changes to 2 files +6 −5

Documentation/git-http-fetch.adoc, http-fetch.c

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 2200f073c4..09b5d675ee 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -50,11 +50,11 @@ commit-id::
 	URL and uses index-pack to generate corresponding .idx and .keep files.
 	The hash is used to determine the name of the temporary file and is
 	arbitrary. The output of index-pack is printed to stdout. Requires
-	--index-pack-args.
+	one or more --index-pack-arg options.
 
---index-pack-args=<args>::
-	For internal use only. The command to run on the contents of the
-	downloaded pack. Arguments are URL-encoded separated by spaces.
+--index-pack-arg=<arg>::
+	For internal use only. An argument to the command run on the contents
+	of the downloaded pack. This option can be specified multiple times.
 
 --recover::
 	Verify that everything reachable from target is fetched.  Used after
diff --git a/http-fetch.c b/http-fetch.c
index f9b6ecb061..601a77c3c1 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)
 
 	if (packfile) {
 		if (!index_pack_args.nr)
-			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
+			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
 
 		fetch_single_packfile(&packfile_hash, argv[arg],
 				      index_pack_args.v);
@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)
 	}
 
 	if (index_pack_args.nr)
-		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
+		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
 
 	if (commits_on_stdin) {
 		commits = walker_targets_stdin(&commit_id, &write_ref);
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 24, 2026, 08:14 UTC in reply to Ted Nyman on lore

[PATCH v4 2/3] http: avoid concurrent appends to partial packs

Pack requests stage downloads in a predictable partial-pack file so an interrupted transfer can be resumed. Both packfile URI and ordinary dumb HTTP requests use this staging path. Opening it in append mode forces each write to the current end of the file, so concurrent responses can append duplicate data and corrupt the pack.

Open the partial pack read-write without O_APPEND and seek once to its current end. Each downloader then retains the offset matching the Range it requested. Because the staging key must uniquely identify immutable pack contents, overlapping responses write the same bytes at the same offsets instead of extending the file with duplicate data.

MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing partial pack exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the staging path. Duplicate that descriptor for index-pack instead of reopening the path after closing the stream; index-pack installs its own pack and the shared staging file is only unlinked, never renamed. Accept HTTP 416 when a partial pack is already complete and let index-pack validate its contents.

Exercise resumed transfers, EOF ranges, overlapping 200 and 206 responses, and unlinking the staging path while index-pack still holds its descriptor. Clarify the staging-key documentation.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc |   5 +-
 http-fetch.c                      |   3 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 250 ++++++++++++++++++++++++++++++
 6 files changed, 295 insertions(+), 25 deletions(-)
Show changes to 6 files +295 −25

Documentation/git-http-fetch.adoc, http-fetch.c, http-push.c, http-walker.c, http.c, t/t5550-http-fetch-dumb.sh

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 09b5d675ee..60ca91cf3a 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -48,8 +48,9 @@ commit-id::
 	line (which is not expected in
 	this case), 'git http-fetch' fetches the packfile directly at the given
 	URL and uses index-pack to generate corresponding .idx and .keep files.
-	The hash is used to determine the name of the temporary file and is
-	arbitrary. The output of index-pack is printed to stdout. Requires
+	The hash is used to determine the name of the temporary file. It need
+	not be the pack hash, but it must uniquely identify the pack contents
+	for resumption. The output of index-pack is printed to stdout. Requires
 	one or more --index-pack-arg options.
 
 --index-pack-arg=<arg>::
diff --git a/http-fetch.c b/http-fetch.c
index 601a77c3c1..05f68f306a 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			struct url_info url;
 			char *nurl = url_normalize(preq->url, &url);
 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
diff --git a/http-push.c b/http-push.c
index 60f6f8f054..ef8abe3908 100644
--- a/http-push.c
+++ b/http-push.c
@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)
 
 	} else if (request->state == RUN_FETCH_PACKED) {
 		int fail = 1;
-		if (request->curl_result != CURLE_OK) {
+		if (request->curl_result != CURLE_OK &&
+		    request->http_code != 416) {
 			fprintf(stderr, "Unable to get pack file %s\n%s",
 				request->url, curl_errorstr);
 		} else {
diff --git a/http-walker.c b/http-walker.c
index b58a3b2a92..abafca84d6 100644
--- a/http-walker.c
+++ b/http-walker.c
@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			error("Unable to get pack file %s\n%s", preq->url,
 			      curl_errorstr);
 			goto abort;
diff --git a/http.c b/http.c
index caccf2108e..a0d399b274 100644
--- a/http.c
+++ b/http.c
@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	int tmpfile_fd;
 	int ret = 0;
 
+	/* Another downloader may unlink the staging path while we index it. */
+	tmpfile_fd = xdup(fileno(preq->packfile));
 	fclose(preq->packfile);
 	preq->packfile = NULL;
-
-	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
+	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
+		die_errno("unable to seek local file %s for pack",
+			  preq->tmpfile.buf);
 
 	ip.git_cmd = 1;
 	ip.in = tmpfile_fd;
@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	else
 		ip.no_stdout = 1;
 
-	if (run_command(&ip)) {
+	if (run_command(&ip))
 		ret = -1;
-		goto cleanup;
-	}
-
-cleanup:
-	close(tmpfile_fd);
 	unlink(preq->tmpfile.buf);
 	return ret;
 }
@@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(
 struct http_pack_request *new_direct_http_pack_request(
 	const unsigned char *packed_git_hash, char *url)
 {
-	off_t prev_posn = 0;
+	off_t prev_posn;
 	struct http_pack_request *preq;
+	int fd;
 
 	CALLOC_ARRAY(preq, 1);
 	strbuf_init(&preq->tmpfile, 0);
-
 	preq->url = url;
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
-	if (!preq->packfile) {
-		error("Unable to open local file %s for pack",
-		      preq->tmpfile.buf);
+	/*
+	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
+	 * existing file; reopen a newly created file so others may unlink it.
+	 */
+	for (;;) {
+		fd = open(preq->tmpfile.buf, O_RDWR);
+		if (fd >= 0 || errno != ENOENT)
+			break;
+		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
+		if (fd >= 0) {
+			close(fd);
+			continue;
+		}
+		if (errno != EEXIST)
+			break;
+	}
+	if (fd < 0) {
+		error_errno("unable to open local file %s for pack",
+			    preq->tmpfile.buf);
 		goto abort;
 	}
+	prev_posn = lseek(fd, 0, SEEK_END);
+	if (prev_posn < 0) {
+		error_errno("unable to seek local file %s for pack",
+			    preq->tmpfile.buf);
+		close(fd);
+		goto abort;
+	}
+	preq->packfile = xfdopen(fd, "w");
 
 	preq->slot = get_active_slot();
 	preq->headers = object_request_headers();
@@ -2762,12 +2783,7 @@ struct http_pack_request *new_direct_http_pack_request(
 	curl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);
 	curl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);
 
-	/*
-	 * If there is data present from a previous transfer attempt,
-	 * resume where it left off
-	 */
-	prev_posn = ftello(preq->packfile);
-	if (prev_posn>0) {
+	if (prev_posn > 0) {
 		if (http_is_verbose)
 			fprintf(stderr,
 				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index f00eeae48f..dcb9667eeb 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -293,6 +293,256 @@ test_expect_success 'http-fetch --packfile' '
 	git -C packfileclient cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile resumes a partial download' '
+	git init packfileclient-resume &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/resume.trace" \
+	git -C packfileclient-resume http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=index-pack --index-pack-arg=--stdin \
+		--index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_grep "Range: bytes=64-" resume.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-resume cat-file -e "$HASH"
+'
+
+test_expect_success 'http-fetch --packfile permits unlink while indexing' '
+	git init packfileclient-unlink &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	write_script git-unlink-index-pack <<-\EOF &&
+	test -f "$GIT_TEST_PACK_TEMP" || exit 1
+	rm "$GIT_TEST_PACK_TEMP" || exit 1
+	exec git index-pack "$@"
+	EOF
+	test_when_finished "rm -f git-unlink-index-pack" &&
+	PATH="$TRASH_DIRECTORY:$PATH" \
+	GIT_TEST_PACK_TEMP="$TRASH_DIRECTORY/$tmpfile" \
+	git -C packfileclient-unlink http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=unlink-index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-unlink cat-file -e "$HASH"
+'
+
+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
+	git init packfileclient-concurrent &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	packhash=$(basename "$p" .pack) &&
+	packhash=${packhash#pack-} &&
+	tmpfile="packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	mkfifo first-ready first-continue &&
+	exec 8<>first-ready &&
+	exec 9<>first-continue &&
+	write_script git-wait-index-pack <<-\EOF &&
+	echo ready >"$GIT_TEST_WAIT_READY" &&
+	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
+	exec git index-pack "$@"
+	EOF
+	{
+		(
+			if ! PATH="$TRASH_DIRECTORY:$PATH" \
+			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
+			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
+			GIT_TRACE_CURL="$TRASH_DIRECTORY/first.trace" \
+			git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+				--index-pack-arg=wait-index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		echo continue >&9
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+		exec 8>&-
+		exec 9>&-
+		rm -f first-ready first-continue git-wait-index-pack
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/second.trace" \
+	git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+		--index-pack-arg=index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
+	echo continue >&9 &&
+	wait "$first_pid" &&
+	printf "pack\t%s\n" "$packhash" >expect &&
+	test_cmp expect first.out &&
+	printf "keep\t%s\n" "$packhash" >expect &&
+	test_cmp expect second.out &&
+	test_grep "Range: bytes=64-" first.trace &&
+	test_grep "Range: bytes=[0-9]*-" second.trace &&
+	test_grep "416 Requested Range Not Satisfiable" second.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-concurrent cat-file -e "$HASH"
+'
+
+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
+	git init packfileclient-overlap &&
+	blob=$(test-tool genrandom pack-overlap 2m |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			hash-object -w --stdin) &&
+	packhash=$(printf "%s\n" "$blob" |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			pack-objects "$TRASH_DIRECTORY/overlap-pack") &&
+	pack="$TRASH_DIRECTORY/overlap-pack-$packhash.pack" &&
+	tmpfile="packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp" &&
+	mkfifo server-ready first-ready &&
+	exec 7<>server-ready &&
+	exec 8<>first-ready &&
+	write_script slow-pack-server "$PERL_PATH" <<-\EOF &&
+	use strict;
+	use warnings;
+	use IO::Socket::INET;
+
+	my ($packfile, $server_ready, $first_ready) = @ARGV;
+	my $completed = 0;
+	END {
+		if (!$completed) {
+			signal_ready($server_ready, "failed");
+			signal_ready($first_ready, "failed");
+		}
+	}
+
+	$SIG{ALRM} = sub { die "timed out serving concurrent pack requests\n" };
+	alarm 60;
+
+	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
+	my $pack = do { local $/; <$in> };
+	close($in) or die "close $packfile: $!";
+	my $server = IO::Socket::INET->new(LocalAddr => "127.0.0.1",
+		LocalPort => 0, Proto => "tcp", Listen => 2, ReuseAddr => 1)
+		or die "listen: $!";
+
+	sub signal_ready {
+		my ($file, $value) = @_;
+		open(my $out, ">", $file) or die "open $file: $!";
+		print $out "$value\n" or die "write $file: $!";
+		close($out) or die "close $file: $!";
+	}
+
+	sub write_all {
+		my ($out, $data) = @_;
+		my $offset = 0;
+		while ($offset < length($data)) {
+			my $written = syswrite($out, $data,
+				length($data) - $offset, $offset);
+			defined($written) && $written or die "write response: $!";
+			$offset += $written;
+		}
+	}
+
+	sub start_response {
+		my $out = $server->accept() or die "accept: $!";
+		<$out> or die "read request: $!";
+		my $start = 0;
+		while (<$out>) {
+			last if /^\r?\n$/;
+			$start = $1 if /^Range: bytes=(\d+)-/i;
+		}
+		$start < length($pack) or die "invalid range $start";
+		my $length = length($pack) - $start;
+		my $middle = int($length / 2);
+		my $status = $start ? "206 Partial Content" : "200 OK";
+		my $headers = "HTTP/1.1 $status\r\n" .
+			"Content-Length: $length\r\n" .
+			($start ? "Content-Range: bytes $start-" .
+				(length($pack) - 1) . "/" . length($pack) . "\r\n" : "") .
+			"Connection: close\r\n\r\n";
+		write_all($out, $headers);
+		write_all($out, substr($pack, $start, $middle));
+		return ($out, $start + $middle);
+	}
+
+	signal_ready($server_ready, $server->sockport());
+	my ($first, $first_pos) = start_response();
+	signal_ready($first_ready, "ready");
+	my ($second, $second_pos) = start_response();
+	write_all($first, substr($pack, $first_pos));
+	write_all($second, substr($pack, $second_pos));
+	close($first) or die "close first response: $!";
+	close($second) or die "close second response: $!";
+	$completed = 1;
+	alarm 0;
+	EOF
+	{
+		"$TRASH_DIRECTORY/slow-pack-server" "$pack" \
+			"$TRASH_DIRECTORY/server-ready" \
+			"$TRASH_DIRECTORY/first-ready" >server.log 2>&1 &
+		server_pid=$!
+	} &&
+	test_when_finished "
+		kill $server_pid 2>/dev/null || :
+		wait $server_pid 2>/dev/null || :
+		exec 7>&-
+		exec 8>&-
+		rm -f server-ready first-ready slow-pack-server
+	" &&
+	read port <&7 &&
+	url="http://127.0.0.1:$port/pack" &&
+	{
+		(
+			if ! GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-first.trace" \
+			GIT_TRACE_CURL_NO_DATA=1 \
+			git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+				--index-pack-arg=index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$url" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	test_path_is_file "$tmpfile" &&
+	test -s "$tmpfile" &&
+	{
+		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
+		GIT_TRACE_CURL_NO_DATA=1 \
+		git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+			--index-pack-arg=index-pack \
+			--index-pack-arg=--stdin --index-pack-arg=--keep \
+			"$url" >second.out &
+		second_pid=$!
+	} &&
+	test_when_finished "
+		kill $second_pid 2>/dev/null || :
+		wait $second_pid 2>/dev/null || :
+	" &&
+	wait "$second_pid" &&
+	wait "$first_pid" &&
+	wait "$server_pid" &&
+	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
+	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
+	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
+	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
+	sort first.out second.out >actual &&
+	test_cmp expect actual &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-overlap cat-file -e "$blob"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 24, 2026, 08:14 UTC in reply to Ted Nyman on lore

[PATCH v4 3/3] fetch-pack: accept "pack" output for packfile URIs

When index-pack finds an existing keep file it reports pack rather than keep. Accept either result from http-fetch, and only register a keep lockfile when this fetch created it.

Read the pack/keep prefix and hash without consuming any following fsck output, validate the reported pack hash against the advertised hash, and exercise a packfile URI fetch with a pre-existing keep file.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 fetch-pack.c           | 33 ++++++++++++++++++---------------
 t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
 2 files changed, 49 insertions(+), 15 deletions(-)
Show changes to 2 files +49 −15

fetch-pack.c, t/t5702-protocol-v2.sh

diff --git a/fetch-pack.c b/fetch-pack.c
index 29c41132ee..e9f24fbd63 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 	}
 
 	for (i = 0; i < packfile_uris.nr; i++) {
+		bool created_keep;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 1];
+		char packhash[GIT_MAX_HEXSZ + 1];
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, 5) < 0 ||
-		    memcmp(packname, "keep\t", 5))
-			die("fetch-pack: expected keep then TAB at start of http-fetch output");
+		if (read_in_full(cmd.out, packhash, 5) != 5 ||
+		    (memcmp(packhash, "keep\t", 5) &&
+		     memcmp(packhash, "pack\t", 5)))
+			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
+		created_keep = !memcmp(packhash, "keep\t", 5);
 
-		if (read_in_full(cmd.out, packname,
-				 the_hash_algo->hexsz + 1) < 0 ||
-		    packname[the_hash_algo->hexsz] != '\n')
-			die("fetch-pack: expected hash then LF at end of http-fetch output");
-
-		packname[the_hash_algo->hexsz] = '\0';
+		if (read_in_full(cmd.out, packhash,
+				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
+		    packhash[the_hash_algo->hexsz] != '\n')
+			die("fetch-pack: expected hash then LF in http-fetch output");
+		packhash[the_hash_algo->hexsz] = '\0';
 
 		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
 
@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (finish_command(&cmd))
 			die("fetch-pack: unable to finish http-fetch");
 
-		if (memcmp(packfile_uris.items[i].string, packname,
+		if (memcmp(packfile_uris.items[i].string, packhash,
 			   the_hash_algo->hexsz))
 			die("fetch-pack: pack downloaded from %s does not match expected hash %.*s",
 			    uri, (int) the_hash_algo->hexsz,
 			    packfile_uris.items[i].string);
 
-		string_list_append_nodup(pack_lockfiles,
-					 xstrfmt("%s/pack/pack-%s.keep",
-						 repo_get_object_directory(the_repository),
-						 packname));
+		if (created_keep)
+			string_list_append_nodup(pack_lockfiles,
+						 xstrfmt("%s/pack/pack-%s.keep",
+							 repo_get_object_directory(the_repository),
+							 packhash));
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);
diff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh
index 74a2b7730b..0f05286de8 100755
--- a/t/t5702-protocol-v2.sh
+++ b/t/t5702-protocol-v2.sh
@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '
 		fetch "$HTTPD_URL/smart/http_parent"
 '
 
+test_expect_success 'packfile URI preserves an existing keep file' '
+	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
+	rm -rf "$P" http_child keep.expect &&
+
+	git init "$P" &&
+	git -C "$P" config uploadpack.allowsidebandall true &&
+
+	echo my-blob >"$P/my-blob" &&
+	git -C "$P" add my-blob &&
+	git -C "$P" commit -m x &&
+	configure_exclusion "$P" my-blob >h &&
+
+	git init http_child &&
+	packhash=$(cat packh) &&
+	keep="http_child/.git/objects/pack/pack-$packhash.keep" &&
+	echo pre-existing >"$keep" &&
+	cp "$keep" keep.expect &&
+
+	GIT_TEST_SIDEBAND_ALL=1 \
+	git -C http_child -c protocol.version=2 \
+		-c fetch.uriprotocols=http,https \
+		fetch "$HTTPD_URL/smart/http_parent" &&
+
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.pack" &&
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.idx" &&
+	test_cmp keep.expect "$keep" &&
+	git -C http_child cat-file -e "$(cat h)"
+'
+
 test_expect_success 'fetching with valid packfile URI but invalid hash fails' '
 	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
 	rm -rf "$P" http_child log &&
-- 
2.55.0.openai.131.g83a728de1eb6
Taylor BlauJul 24, 2026, 21:38 UTC in reply to Ted Nyman on lore

Re: [PATCH v4 1/3] http-fetch: correct --index-pack-arg documentation

On Fri, Jul 24, 2026 at 01:14:23AM -0700, Ted Nyman wrote:
> The --packfile mode accepts one --index-pack-arg=<arg> option per
> argument passed to index-pack, but its documentation and option
> dependency errors still refer to the plural --index-pack-args form.

Good find, it looks like this dates all the way back to 27e35ba6c6 (http-fetch: allow custom index-pack args, 2021-02-22). Thanks for taking the time to correct it.

Show 17 quoted lines
> diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
> index 2200f073c4..09b5d675ee 100644
> --- a/Documentation/git-http-fetch.adoc
> +++ b/Documentation/git-http-fetch.adoc
> @@ -50,11 +50,11 @@ commit-id::
>  	URL and uses index-pack to generate corresponding .idx and .keep files.
>  	The hash is used to determine the name of the temporary file and is
>  	arbitrary. The output of index-pack is printed to stdout. Requires
> -	--index-pack-args.
> +	one or more --index-pack-arg options.
>
> ---index-pack-args=<args>::
> -	For internal use only. The command to run on the contents of the
> -	downloaded pack. Arguments are URL-encoded separated by spaces.
> +--index-pack-arg=<arg>::
> +	For internal use only. An argument to the command run on the contents
> +	of the downloaded pack. This option can be specified multiple times.

Interesting. The plural "--index-pack-args" form says that it specifies the command to run on the downloaded pack, as well as arguments which are separated by spaces. Two thoughts:

 - I think the "arguments are URL-encoded separated by spaces" claim was
   not true even in 27e35ba6c6, so dropping that seems like a strict
   improvement to me.
 - The new form says "An argument to the command run on [...]", but I
   believe that this option is also used to specify the name of the
   command to run itself. I wonder if it may be worth saying something
   like "The first instance specifies the command to run. Subsequent
   occurrences specify its arguments."
> diff --git a/http-fetch.c b/http-fetch.c
> index f9b6ecb061..601a77c3c1 100644
> --- a/http-fetch.c
> +++ b/http-fetch.c

Changes in this file look reasonable. Likewise, it makes sense that we do not have any changes in the test suite, since this option did not exist in a plural in the first place ;-).

Thanks, Taylor

Taylor BlauJul 24, 2026, 21:46 UTC in reply to Ted Nyman on lore

Re: [PATCH v4 3/3] fetch-pack: accept "pack" output for packfile URIs

On Fri, Jul 24, 2026 at 01:14:25AM -0700, Ted Nyman wrote:
Show 27 quoted lines
> When index-pack finds an existing keep file it reports pack rather than
> keep. Accept either result from http-fetch, and only register a keep
> lockfile when this fetch created it.
>
> Read the pack/keep prefix and hash without consuming any following fsck
> output, validate the reported pack hash against the advertised hash, and
> exercise a packfile URI fetch with a pre-existing keep file.
>
> Signed-off-by: Ted Nyman <tnyman@openai.com>
> ---
>  fetch-pack.c           | 33 ++++++++++++++++++---------------
>  t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
>  2 files changed, 49 insertions(+), 15 deletions(-)
>
> diff --git a/fetch-pack.c b/fetch-pack.c
> index 29c41132ee..e9f24fbd63 100644
> --- a/fetch-pack.c
> +++ b/fetch-pack.c
> @@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
>  	}
>
>  	for (i = 0; i < packfile_uris.nr; i++) {
> +		bool created_keep;
>  		int j;
>  		struct child_process cmd = CHILD_PROCESS_INIT;
> -		char packname[GIT_MAX_HEXSZ + 1];
> +		char packhash[GIT_MAX_HEXSZ + 1];

OK, so we keep track of whether or not we got "keep" as part of the output.

While here, "packname" is renamed to "packhash", which I think is reasonable, especially to indicate that the buffer is sized accordingly. We happen to read the preceding "pack" or "keep" into that same buffer, which I think is fine. If we wanted to be pedantic we could read that into a separate buffer, but I don't think such separation is necessary.

Show 15 quoted lines
>  		const char *uri = packfile_uris.items[i].string +
>  			the_hash_algo->hexsz + 1;
>
> @@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
>  		if (start_command(&cmd))
>  			die("fetch-pack: unable to spawn http-fetch");
>
> -		if (read_in_full(cmd.out, packname, 5) < 0 ||
> -		    memcmp(packname, "keep\t", 5))
> -			die("fetch-pack: expected keep then TAB at start of http-fetch output");
> +		if (read_in_full(cmd.out, packhash, 5) != 5 ||
> +		    (memcmp(packhash, "keep\t", 5) &&
> +		     memcmp(packhash, "pack\t", 5)))
> +			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
> +		created_keep = !memcmp(packhash, "keep\t", 5);
Makes sense.
Show 12 quoted lines
>
> -		if (read_in_full(cmd.out, packname,
> -				 the_hash_algo->hexsz + 1) < 0 ||
> -		    packname[the_hash_algo->hexsz] != '\n')
> -			die("fetch-pack: expected hash then LF at end of http-fetch output");
> -
> -		packname[the_hash_algo->hexsz] = '\0';
> +		if (read_in_full(cmd.out, packhash,
> +				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
> +		    packhash[the_hash_algo->hexsz] != '\n')
> +			die("fetch-pack: expected hash then LF in http-fetch output");
> +		packhash[the_hash_algo->hexsz] = '\0';
Likewise. The rest of this file and the test also look good.

Thanks, Taylor

Jeff KingJul 25, 2026, 09:09 UTC in reply to Junio C Hamano on lore

Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads

On Thu, Jul 23, 2026 at 09:43:14PM -0700, Junio C Hamano wrote:
Show 5 quoted lines
> When merged into 'seen', this topic seems to cause t5550 to hang
> fairly consistently.  It is not surprising, considering that the
> topic adds roughly 240 lines to the test script in question.  It is
> entirely possible that we are seeing an existing breakage from
> another topic in 'seen' that is exposed by the additional tests.

I didn't get any hang locally, but running t5550 with --stress causes around half of the runs to fail immediately. That continues to be true with v4. So there is presumably some race condition still present.

The failing test is the big one (34) and the failing command is the "test -s $tmpfile" call. It looks like the pack has already been indexed (at least by the time I look at the on-disk state of a failed example).

I don't immediately see the issue, though. I could believe that extra load fakes out any sleep-based timing tricks, but it looks like the test tries to use FIFOs to do everything deterministically.

Diffing the overlap-first.trace file between a working case and a failing one, I see (skipping past uninteresting port differences) this hunk at the end:

  @@ -17,4 +17,5 @@
   <= Recv header: Connection: close
   <= Recv header, 0000000002 bytes (0x00000002)
   <= Recv header:
  -== Info: shutting down connection #0
  +== Info: end of response with 1048917 bytes missing
  +== Info: closing connection #0

So curl sees a hangup on the first connection (even though the second one hasn't even started yet!). I'm not sure why, though. There's nothing useful in the server.log file. I tried stracing the server process but it didn't show much of interest. Both cases write "ready" to first-ready, and then the success case immediately sees an accept() for the second connection. The failing case waits in accept() and then eventually calls SIGALRM (which is way after the failure happens; the test has already bailed and so the second connection never comes in).

So from the perspective of the server process, everything is fine, but curl complains that it didn't get all of the bytes. Weird. The strace shows both writing the first 1MB as expected. It's like the connection gets hung up for some reason, but I can't tell why or by whom.

-Peff
Jeff KingJul 25, 2026, 09:21 UTC in reply to Jeff King on lore

Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads

On Sat, Jul 25, 2026 at 05:09:11AM -0400, Jeff King wrote:
> So from the perspective of the server process, everything is fine, but
> curl complains that it didn't get all of the bytes. Weird. The strace
> shows both writing the first 1MB as expected. It's like the connection
> gets hung up for some reason, but I can't tell why or by whom.

Hmph. I tried stracing on the client side, and we indeed see an EOF on the socket:

  recvfrom(12, "", 16384, 0, NULL, NULL) = 0

I'm really puzzled why that is the case, though, as the server side did not close() or exit.

I do wonder if it would be possible to write this script so that it triggers via apache, like the rest of our http tests. Doing our own socket handling here is a potential source of bugs.

-Peff
Jeff KingJul 25, 2026, 10:02 UTC in reply to Jeff King on lore

Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads

On Sat, Jul 25, 2026 at 05:21:54AM -0400, Jeff King wrote:
Show 14 quoted lines
> On Sat, Jul 25, 2026 at 05:09:11AM -0400, Jeff King wrote:
> 
> > So from the perspective of the server process, everything is fine, but
> > curl complains that it didn't get all of the bytes. Weird. The strace
> > shows both writing the first 1MB as expected. It's like the connection
> > gets hung up for some reason, but I can't tell why or by whom.
> 
> Hmph. I tried stracing on the client side, and we indeed see an EOF
> on the socket:
> 
>   recvfrom(12, "", 16384, 0, NULL, NULL) = 0
> 
> I'm really puzzled why that is the case, though, as the server side did
> not close() or exit.

So I'm still puzzled by all of this, but I think it's mostly a red herring with respect to the actual "test -s" race.

The server tells us when it has written the first 1MB to the client, and then we check that "test -s" is showing something in the on-disk tempfile we're downloading. But there's no guarantee that just because the server called write() that client has yet received the data, let alone written it to disk.

I can't think of a synchronization point we could use here. We're waiting on curl to have passed the bytes to fwrite() and for it to have actually synced to disk. We either have to poll or modify http.c to write "yes, we got some bytes!" to a fifo. Both are pretty gross.

I wonder if we could just drop that "test -s" entirely. We'd _usually_ see some bytes written before the second request starts. But it's OK if we don't. It just means the test is working in the reverse order (the second request may write its bytes first, and then the first one is the one "overwriting" it). I.e., the two are symmetric from our perspective.

-Peff
Jeff KingJul 25, 2026, 10:10 UTC in reply to Jeff King on lore

Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads

On Sat, Jul 25, 2026 at 06:02:51AM -0400, Jeff King wrote:
Show 5 quoted lines
> I wonder if we could just drop that "test -s" entirely. We'd _usually_
> see some bytes written before the second request starts. But it's OK if
> we don't. It just means the test is working in the reverse order (the
> second request may write its bytes first, and then the first one is the
> one "overwriting" it). I.e., the two are symmetric from our perspective.
Yeah, doing this:
Show changes to t/t5550-http-fetch-dumb.sh +0 −5
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index dcb9667eeb..07aa218049 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -516,7 +516,6 @@ test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt a
 	read ready <&8 &&
 	test "$ready" = ready &&
 	test_path_is_file "$tmpfile" &&
-	test -s "$tmpfile" &&
 	{
 		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
 		GIT_TRACE_CURL_NO_DATA=1 \
@@ -533,9 +532,6 @@ test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt a
 	wait "$second_pid" &&
 	wait "$first_pid" &&
 	wait "$server_pid" &&
-	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
-	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
-	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
 	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
 	sort first.out second.out >actual &&
 	test_cmp expect actual &&

is enough to make it pass reliably under --stress for me. We have to
drop the trace greps, because we don't actually know whether each
request will use a range or not. We'd _usually_ see a range for the
second one, but it's possible it might still see a zero-byte file. I
guess we probably see a "200" reliably for the first request, but it's
not all that interesting.

We can leave the test_path_is_file check, because we open the file
before making the request (it is only the actual writing of bytes that
is racy).

-Peff
Junio C HamanoJul 25, 2026, 16:20 UTC in reply to Jeff King on lore

Re: [PATCH v3 0/3] packfile URIs: support concurrent downloads

Jeff King <peff@peff.net> writes:
Show 10 quoted lines
> I can't think of a synchronization point we could use here. We're
> waiting on curl to have passed the bytes to fwrite() and for it to have
> actually synced to disk. We either have to poll or modify http.c to
> write "yes, we got some bytes!" to a fifo. Both are pretty gross.
>
> I wonder if we could just drop that "test -s" entirely. We'd _usually_
> see some bytes written before the second request starts. But it's OK if
> we don't. It just means the test is working in the reverse order (the
> second request may write its bytes first, and then the first one is the
> one "overwriting" it). I.e., the two are symmetric from our perspective.

Yeah, that sounds quite sensible. Thanks for digging.

Ted NymanJul 26, 2026, 06:44 UTC in reply to Ted Nyman on lore

[PATCH v5 0/3] packfile URIs: support concurrent downloads

Packfile URI and dumb HTTP downloads stage packs at objects/pack/pack-<hash>.pack.temp so an interrupted transfer can resume. Opening that file in append mode forces every write to its current end. Two Git processes fetching the same pack into one object database can therefore append duplicate data and corrupt the pack.

The first patch separates the unrelated --index-pack-arg documentation and error-message correction requested during review.

The second patch keeps the predictable staging name but removes append mode. Each downloader seeks once to the current end, requests the corresponding Range, and writes using its own descriptor offset. Since the staging key must identify immutable pack contents, overlapping responses write identical bytes at identical offsets. There is no need for pwrite(2) or cross-process coordination, and resumption continues to work for both packfile URI and ordinary dumb HTTP downloads.

A downloader can also find that the partial pack has completed and request a range starting at EOF. Servers may respond with HTTP 416 in that case. Treat the response as a completed download and let index-pack validate the pack.

On MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing staging file exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the path. Keep the open descriptor for index-pack; it installs its own pack, so the shared staging file is only unlinked, never renamed.

The third patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process. Read only the prefix and hash so any following fsck output remains available to fetch-pack.

The tests cover resumption, a completed partial returning 416, overlapping downloads, unlinking the staging path while index-pack holds its descriptor, and a pre-existing .keep file. The unlink test does not require FIFOs, so it can exercise MinGW's sharing behavior even though the concurrent-download tests are skipped there.

Changes since v4:
  * Clarify that the first --index-pack-arg specifies the command and
    subsequent instances specify its arguments.
  * Drop assumptions about which concurrent response reaches the
    staging file first. Either write order exercises the same
    overlapping-download behavior.
  * No production code changes.

The overlapping-download test passes 240 runs with 12 parallel stress jobs.

The v4 discussion is at:
  https://lore.kernel.org/git/cover.1784874850.git.tnyman@openai.com/
Ted Nyman (3):
  http-fetch: correct --index-pack-arg documentation
  http: avoid concurrent appends to partial packs
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  14 +-
 fetch-pack.c                      |  33 ++--
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 246 ++++++++++++++++++++++++++++++
 t/t5702-protocol-v2.sh            |  31 ++++
 8 files changed, 347 insertions(+), 46 deletions(-)
Range-diff against v4:
1:  a6a40b8046 ! 1:  a79af009ea http-fetch: correct --index-pack-arg documentation
    @@ Documentation/git-http-fetch.adoc: commit-id::
     -	For internal use only. The command to run on the contents of the
     -	downloaded pack. Arguments are URL-encoded separated by spaces.
     +--index-pack-arg=<arg>::
    -+	For internal use only. An argument to the command run on the contents
    -+	of the downloaded pack. This option can be specified multiple times.
    ++	For internal use only. The first instance specifies the command run on
    ++	the contents of the downloaded pack. Subsequent instances specify its
    ++	arguments.
      
      --recover::
      	Verify that everything reachable from target is fetched.  Used after
2:  144c98cdfa ! 2:  d9667c93b0 http: avoid concurrent appends to partial packs
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	read ready <&8 &&
     +	test "$ready" = ready &&
     +	test_path_is_file "$tmpfile" &&
    -+	test -s "$tmpfile" &&
     +	{
     +		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
     +		GIT_TRACE_CURL_NO_DATA=1 \
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	wait "$second_pid" &&
     +	wait "$first_pid" &&
     +	wait "$server_pid" &&
    -+	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
    -+	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
    -+	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
     +	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
     +	sort first.out second.out >actual &&
     +	test_cmp expect actual &&
3:  d9063deb60 = 3:  fee6f292cb fetch-pack: accept "pack" output for packfile URIs
base-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 26, 2026, 06:44 UTC in reply to Ted Nyman on lore

[PATCH v5 1/3] http-fetch: correct --index-pack-arg documentation

The --packfile mode accepts one --index-pack-arg=<arg> option per argument passed to index-pack, but its documentation and option dependency errors still refer to the plural --index-pack-args form.

Correct the spelling and describe the repeatable per-argument form.
Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc | 9 +++++----
 http-fetch.c                      | 4 ++--
 2 files changed, 7 insertions(+), 6 deletions(-)
Show changes to 2 files +7 −5

Documentation/git-http-fetch.adoc, http-fetch.c

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 2200f073c4..12036e65e9 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -50,11 +50,12 @@ commit-id::
 	URL and uses index-pack to generate corresponding .idx and .keep files.
 	The hash is used to determine the name of the temporary file and is
 	arbitrary. The output of index-pack is printed to stdout. Requires
-	--index-pack-args.
+	one or more --index-pack-arg options.
 
---index-pack-args=<args>::
-	For internal use only. The command to run on the contents of the
-	downloaded pack. Arguments are URL-encoded separated by spaces.
+--index-pack-arg=<arg>::
+	For internal use only. The first instance specifies the command run on
+	the contents of the downloaded pack. Subsequent instances specify its
+	arguments.
 
 --recover::
 	Verify that everything reachable from target is fetched.  Used after
diff --git a/http-fetch.c b/http-fetch.c
index f9b6ecb061..601a77c3c1 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)
 
 	if (packfile) {
 		if (!index_pack_args.nr)
-			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
+			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
 
 		fetch_single_packfile(&packfile_hash, argv[arg],
 				      index_pack_args.v);
@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)
 	}
 
 	if (index_pack_args.nr)
-		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
+		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
 
 	if (commits_on_stdin) {
 		commits = walker_targets_stdin(&commit_id, &write_ref);
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 26, 2026, 06:44 UTC in reply to Ted Nyman on lore

[PATCH v5 2/3] http: avoid concurrent appends to partial packs

Pack requests stage downloads in a predictable partial-pack file so an interrupted transfer can be resumed. Both packfile URI and ordinary dumb HTTP requests use this staging path. Opening it in append mode forces each write to the current end of the file, so concurrent responses can append duplicate data and corrupt the pack.

Open the partial pack read-write without O_APPEND and seek once to its current end. Each downloader then retains the offset matching the Range it requested. Because the staging key must uniquely identify immutable pack contents, overlapping responses write the same bytes at the same offsets instead of extending the file with duplicate data.

MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing partial pack exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the staging path. Duplicate that descriptor for index-pack instead of reopening the path after closing the stream; index-pack installs its own pack and the shared staging file is only unlinked, never renamed. Accept HTTP 416 when a partial pack is already complete and let index-pack validate its contents.

Exercise resumed transfers, EOF ranges, overlapping 200 and 206 responses, and unlinking the staging path while index-pack still holds its descriptor. Clarify the staging-key documentation.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc |   5 +-
 http-fetch.c                      |   3 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 246 ++++++++++++++++++++++++++++++
 6 files changed, 291 insertions(+), 25 deletions(-)
Show changes to 6 files +291 −25

Documentation/git-http-fetch.adoc, http-fetch.c, http-push.c, http-walker.c, http.c, t/t5550-http-fetch-dumb.sh

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 12036e65e9..45e0d3d07c 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -48,8 +48,9 @@ commit-id::
 	line (which is not expected in
 	this case), 'git http-fetch' fetches the packfile directly at the given
 	URL and uses index-pack to generate corresponding .idx and .keep files.
-	The hash is used to determine the name of the temporary file and is
-	arbitrary. The output of index-pack is printed to stdout. Requires
+	The hash is used to determine the name of the temporary file. It need
+	not be the pack hash, but it must uniquely identify the pack contents
+	for resumption. The output of index-pack is printed to stdout. Requires
 	one or more --index-pack-arg options.
 
 --index-pack-arg=<arg>::
diff --git a/http-fetch.c b/http-fetch.c
index 601a77c3c1..05f68f306a 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			struct url_info url;
 			char *nurl = url_normalize(preq->url, &url);
 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
diff --git a/http-push.c b/http-push.c
index 60f6f8f054..ef8abe3908 100644
--- a/http-push.c
+++ b/http-push.c
@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)
 
 	} else if (request->state == RUN_FETCH_PACKED) {
 		int fail = 1;
-		if (request->curl_result != CURLE_OK) {
+		if (request->curl_result != CURLE_OK &&
+		    request->http_code != 416) {
 			fprintf(stderr, "Unable to get pack file %s\n%s",
 				request->url, curl_errorstr);
 		} else {
diff --git a/http-walker.c b/http-walker.c
index b58a3b2a92..abafca84d6 100644
--- a/http-walker.c
+++ b/http-walker.c
@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			error("Unable to get pack file %s\n%s", preq->url,
 			      curl_errorstr);
 			goto abort;
diff --git a/http.c b/http.c
index caccf2108e..a0d399b274 100644
--- a/http.c
+++ b/http.c
@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	int tmpfile_fd;
 	int ret = 0;
 
+	/* Another downloader may unlink the staging path while we index it. */
+	tmpfile_fd = xdup(fileno(preq->packfile));
 	fclose(preq->packfile);
 	preq->packfile = NULL;
-
-	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
+	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
+		die_errno("unable to seek local file %s for pack",
+			  preq->tmpfile.buf);
 
 	ip.git_cmd = 1;
 	ip.in = tmpfile_fd;
@@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	else
 		ip.no_stdout = 1;
 
-	if (run_command(&ip)) {
+	if (run_command(&ip))
 		ret = -1;
-		goto cleanup;
-	}
-
-cleanup:
-	close(tmpfile_fd);
 	unlink(preq->tmpfile.buf);
 	return ret;
 }
@@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(
 struct http_pack_request *new_direct_http_pack_request(
 	const unsigned char *packed_git_hash, char *url)
 {
-	off_t prev_posn = 0;
+	off_t prev_posn;
 	struct http_pack_request *preq;
+	int fd;
 
 	CALLOC_ARRAY(preq, 1);
 	strbuf_init(&preq->tmpfile, 0);
-
 	preq->url = url;
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
-	if (!preq->packfile) {
-		error("Unable to open local file %s for pack",
-		      preq->tmpfile.buf);
+	/*
+	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
+	 * existing file; reopen a newly created file so others may unlink it.
+	 */
+	for (;;) {
+		fd = open(preq->tmpfile.buf, O_RDWR);
+		if (fd >= 0 || errno != ENOENT)
+			break;
+		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
+		if (fd >= 0) {
+			close(fd);
+			continue;
+		}
+		if (errno != EEXIST)
+			break;
+	}
+	if (fd < 0) {
+		error_errno("unable to open local file %s for pack",
+			    preq->tmpfile.buf);
 		goto abort;
 	}
+	prev_posn = lseek(fd, 0, SEEK_END);
+	if (prev_posn < 0) {
+		error_errno("unable to seek local file %s for pack",
+			    preq->tmpfile.buf);
+		close(fd);
+		goto abort;
+	}
+	preq->packfile = xfdopen(fd, "w");
 
 	preq->slot = get_active_slot();
 	preq->headers = object_request_headers();
@@ -2762,12 +2783,7 @@ struct http_pack_request *new_direct_http_pack_request(
 	curl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);
 	curl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);
 
-	/*
-	 * If there is data present from a previous transfer attempt,
-	 * resume where it left off
-	 */
-	prev_posn = ftello(preq->packfile);
-	if (prev_posn>0) {
+	if (prev_posn > 0) {
 		if (http_is_verbose)
 			fprintf(stderr,
 				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index f00eeae48f..07aa218049 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -293,6 +293,252 @@ test_expect_success 'http-fetch --packfile' '
 	git -C packfileclient cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile resumes a partial download' '
+	git init packfileclient-resume &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/resume.trace" \
+	git -C packfileclient-resume http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=index-pack --index-pack-arg=--stdin \
+		--index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_grep "Range: bytes=64-" resume.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-resume cat-file -e "$HASH"
+'
+
+test_expect_success 'http-fetch --packfile permits unlink while indexing' '
+	git init packfileclient-unlink &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	write_script git-unlink-index-pack <<-\EOF &&
+	test -f "$GIT_TEST_PACK_TEMP" || exit 1
+	rm "$GIT_TEST_PACK_TEMP" || exit 1
+	exec git index-pack "$@"
+	EOF
+	test_when_finished "rm -f git-unlink-index-pack" &&
+	PATH="$TRASH_DIRECTORY:$PATH" \
+	GIT_TEST_PACK_TEMP="$TRASH_DIRECTORY/$tmpfile" \
+	git -C packfileclient-unlink http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=unlink-index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-unlink cat-file -e "$HASH"
+'
+
+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
+	git init packfileclient-concurrent &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	packhash=$(basename "$p" .pack) &&
+	packhash=${packhash#pack-} &&
+	tmpfile="packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	mkfifo first-ready first-continue &&
+	exec 8<>first-ready &&
+	exec 9<>first-continue &&
+	write_script git-wait-index-pack <<-\EOF &&
+	echo ready >"$GIT_TEST_WAIT_READY" &&
+	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
+	exec git index-pack "$@"
+	EOF
+	{
+		(
+			if ! PATH="$TRASH_DIRECTORY:$PATH" \
+			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
+			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
+			GIT_TRACE_CURL="$TRASH_DIRECTORY/first.trace" \
+			git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+				--index-pack-arg=wait-index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		echo continue >&9
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+		exec 8>&-
+		exec 9>&-
+		rm -f first-ready first-continue git-wait-index-pack
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/second.trace" \
+	git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
+		--index-pack-arg=index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
+	echo continue >&9 &&
+	wait "$first_pid" &&
+	printf "pack\t%s\n" "$packhash" >expect &&
+	test_cmp expect first.out &&
+	printf "keep\t%s\n" "$packhash" >expect &&
+	test_cmp expect second.out &&
+	test_grep "Range: bytes=64-" first.trace &&
+	test_grep "Range: bytes=[0-9]*-" second.trace &&
+	test_grep "416 Requested Range Not Satisfiable" second.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-concurrent cat-file -e "$HASH"
+'
+
+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
+	git init packfileclient-overlap &&
+	blob=$(test-tool genrandom pack-overlap 2m |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			hash-object -w --stdin) &&
+	packhash=$(printf "%s\n" "$blob" |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			pack-objects "$TRASH_DIRECTORY/overlap-pack") &&
+	pack="$TRASH_DIRECTORY/overlap-pack-$packhash.pack" &&
+	tmpfile="packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp" &&
+	mkfifo server-ready first-ready &&
+	exec 7<>server-ready &&
+	exec 8<>first-ready &&
+	write_script slow-pack-server "$PERL_PATH" <<-\EOF &&
+	use strict;
+	use warnings;
+	use IO::Socket::INET;
+
+	my ($packfile, $server_ready, $first_ready) = @ARGV;
+	my $completed = 0;
+	END {
+		if (!$completed) {
+			signal_ready($server_ready, "failed");
+			signal_ready($first_ready, "failed");
+		}
+	}
+
+	$SIG{ALRM} = sub { die "timed out serving concurrent pack requests\n" };
+	alarm 60;
+
+	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
+	my $pack = do { local $/; <$in> };
+	close($in) or die "close $packfile: $!";
+	my $server = IO::Socket::INET->new(LocalAddr => "127.0.0.1",
+		LocalPort => 0, Proto => "tcp", Listen => 2, ReuseAddr => 1)
+		or die "listen: $!";
+
+	sub signal_ready {
+		my ($file, $value) = @_;
+		open(my $out, ">", $file) or die "open $file: $!";
+		print $out "$value\n" or die "write $file: $!";
+		close($out) or die "close $file: $!";
+	}
+
+	sub write_all {
+		my ($out, $data) = @_;
+		my $offset = 0;
+		while ($offset < length($data)) {
+			my $written = syswrite($out, $data,
+				length($data) - $offset, $offset);
+			defined($written) && $written or die "write response: $!";
+			$offset += $written;
+		}
+	}
+
+	sub start_response {
+		my $out = $server->accept() or die "accept: $!";
+		<$out> or die "read request: $!";
+		my $start = 0;
+		while (<$out>) {
+			last if /^\r?\n$/;
+			$start = $1 if /^Range: bytes=(\d+)-/i;
+		}
+		$start < length($pack) or die "invalid range $start";
+		my $length = length($pack) - $start;
+		my $middle = int($length / 2);
+		my $status = $start ? "206 Partial Content" : "200 OK";
+		my $headers = "HTTP/1.1 $status\r\n" .
+			"Content-Length: $length\r\n" .
+			($start ? "Content-Range: bytes $start-" .
+				(length($pack) - 1) . "/" . length($pack) . "\r\n" : "") .
+			"Connection: close\r\n\r\n";
+		write_all($out, $headers);
+		write_all($out, substr($pack, $start, $middle));
+		return ($out, $start + $middle);
+	}
+
+	signal_ready($server_ready, $server->sockport());
+	my ($first, $first_pos) = start_response();
+	signal_ready($first_ready, "ready");
+	my ($second, $second_pos) = start_response();
+	write_all($first, substr($pack, $first_pos));
+	write_all($second, substr($pack, $second_pos));
+	close($first) or die "close first response: $!";
+	close($second) or die "close second response: $!";
+	$completed = 1;
+	alarm 0;
+	EOF
+	{
+		"$TRASH_DIRECTORY/slow-pack-server" "$pack" \
+			"$TRASH_DIRECTORY/server-ready" \
+			"$TRASH_DIRECTORY/first-ready" >server.log 2>&1 &
+		server_pid=$!
+	} &&
+	test_when_finished "
+		kill $server_pid 2>/dev/null || :
+		wait $server_pid 2>/dev/null || :
+		exec 7>&-
+		exec 8>&-
+		rm -f server-ready first-ready slow-pack-server
+	" &&
+	read port <&7 &&
+	url="http://127.0.0.1:$port/pack" &&
+	{
+		(
+			if ! GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-first.trace" \
+			GIT_TRACE_CURL_NO_DATA=1 \
+			git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+				--index-pack-arg=index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$url" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	test_path_is_file "$tmpfile" &&
+	{
+		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
+		GIT_TRACE_CURL_NO_DATA=1 \
+		git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+			--index-pack-arg=index-pack \
+			--index-pack-arg=--stdin --index-pack-arg=--keep \
+			"$url" >second.out &
+		second_pid=$!
+	} &&
+	test_when_finished "
+		kill $second_pid 2>/dev/null || :
+		wait $second_pid 2>/dev/null || :
+	" &&
+	wait "$second_pid" &&
+	wait "$first_pid" &&
+	wait "$server_pid" &&
+	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
+	sort first.out second.out >actual &&
+	test_cmp expect actual &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-overlap cat-file -e "$blob"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 26, 2026, 06:44 UTC in reply to Ted Nyman on lore

[PATCH v5 3/3] fetch-pack: accept "pack" output for packfile URIs

When index-pack finds an existing keep file it reports pack rather than keep. Accept either result from http-fetch, and only register a keep lockfile when this fetch created it.

Read the pack/keep prefix and hash without consuming any following fsck output, validate the reported pack hash against the advertised hash, and exercise a packfile URI fetch with a pre-existing keep file.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 fetch-pack.c           | 33 ++++++++++++++++++---------------
 t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
 2 files changed, 49 insertions(+), 15 deletions(-)
Show changes to 2 files +49 −15

fetch-pack.c, t/t5702-protocol-v2.sh

diff --git a/fetch-pack.c b/fetch-pack.c
index 29c41132ee..e9f24fbd63 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 	}
 
 	for (i = 0; i < packfile_uris.nr; i++) {
+		bool created_keep;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 1];
+		char packhash[GIT_MAX_HEXSZ + 1];
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, 5) < 0 ||
-		    memcmp(packname, "keep\t", 5))
-			die("fetch-pack: expected keep then TAB at start of http-fetch output");
+		if (read_in_full(cmd.out, packhash, 5) != 5 ||
+		    (memcmp(packhash, "keep\t", 5) &&
+		     memcmp(packhash, "pack\t", 5)))
+			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
+		created_keep = !memcmp(packhash, "keep\t", 5);
 
-		if (read_in_full(cmd.out, packname,
-				 the_hash_algo->hexsz + 1) < 0 ||
-		    packname[the_hash_algo->hexsz] != '\n')
-			die("fetch-pack: expected hash then LF at end of http-fetch output");
-
-		packname[the_hash_algo->hexsz] = '\0';
+		if (read_in_full(cmd.out, packhash,
+				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
+		    packhash[the_hash_algo->hexsz] != '\n')
+			die("fetch-pack: expected hash then LF in http-fetch output");
+		packhash[the_hash_algo->hexsz] = '\0';
 
 		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
 
@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (finish_command(&cmd))
 			die("fetch-pack: unable to finish http-fetch");
 
-		if (memcmp(packfile_uris.items[i].string, packname,
+		if (memcmp(packfile_uris.items[i].string, packhash,
 			   the_hash_algo->hexsz))
 			die("fetch-pack: pack downloaded from %s does not match expected hash %.*s",
 			    uri, (int) the_hash_algo->hexsz,
 			    packfile_uris.items[i].string);
 
-		string_list_append_nodup(pack_lockfiles,
-					 xstrfmt("%s/pack/pack-%s.keep",
-						 repo_get_object_directory(the_repository),
-						 packname));
+		if (created_keep)
+			string_list_append_nodup(pack_lockfiles,
+						 xstrfmt("%s/pack/pack-%s.keep",
+							 repo_get_object_directory(the_repository),
+							 packhash));
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);
diff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh
index 74a2b7730b..0f05286de8 100755
--- a/t/t5702-protocol-v2.sh
+++ b/t/t5702-protocol-v2.sh
@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '
 		fetch "$HTTPD_URL/smart/http_parent"
 '
 
+test_expect_success 'packfile URI preserves an existing keep file' '
+	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
+	rm -rf "$P" http_child keep.expect &&
+
+	git init "$P" &&
+	git -C "$P" config uploadpack.allowsidebandall true &&
+
+	echo my-blob >"$P/my-blob" &&
+	git -C "$P" add my-blob &&
+	git -C "$P" commit -m x &&
+	configure_exclusion "$P" my-blob >h &&
+
+	git init http_child &&
+	packhash=$(cat packh) &&
+	keep="http_child/.git/objects/pack/pack-$packhash.keep" &&
+	echo pre-existing >"$keep" &&
+	cp "$keep" keep.expect &&
+
+	GIT_TEST_SIDEBAND_ALL=1 \
+	git -C http_child -c protocol.version=2 \
+		-c fetch.uriprotocols=http,https \
+		fetch "$HTTPD_URL/smart/http_parent" &&
+
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.pack" &&
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.idx" &&
+	test_cmp keep.expect "$keep" &&
+	git -C http_child cat-file -e "$(cat h)"
+'
+
 test_expect_success 'fetching with valid packfile URI but invalid hash fails' '
 	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
 	rm -rf "$P" http_child log &&
-- 
2.55.0.openai.131.g83a728de1eb6
Jeff KingJul 26, 2026, 09:20 UTC in reply to Ted Nyman on lore

Re: [PATCH v5 2/3] http: avoid concurrent appends to partial packs

On Sat, Jul 25, 2026 at 11:44:47PM -0700, Ted Nyman wrote:
Show 11 quoted lines
> Pack requests stage downloads in a predictable partial-pack file so an
> interrupted transfer can be resumed. Both packfile URI and ordinary dumb
> HTTP requests use this staging path. Opening it in append mode forces
> each write to the current end of the file, so concurrent responses can
> append duplicate data and corrupt the pack.
> 
> Open the partial pack read-write without O_APPEND and seek once to its
> current end. Each downloader then retains the offset matching the Range
> it requested. Because the staging key must uniquely identify immutable
> pack contents, overlapping responses write the same bytes at the same
> offsets instead of extending the file with duplicate data.

OK. I still think this is kind of horrible and gross, but I can't think of a reason it won't work (at least on POSIX-ish systems) and it solves the problem with minimal changes and risk of regression.

I wondered about racing with another concurrent writer on the seek, but I think it is OK. We seek immediately, and then use that offset (which we get from another seek, replacing ftell()) as the value for our range request. So even if somebody else advances the file, we have _some_ atomic value that we'll start writing to ourselves, and the worst case is redundantly requesting a few bytes.

Show 7 quoted lines
> MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
> existing file. Create a missing partial pack exclusively, close it, and
> reopen it without O_CREAT so every retained descriptor permits another
> downloader to unlink the staging path. Duplicate that descriptor for
> index-pack instead of reopening the path after closing the stream;
> index-pack installs its own pack and the shared staging file is only
> unlinked, never renamed.

This part I have no real knowledge or opinion on the Windows bits (or if there's an easier way to do it).

> Accept HTTP 416 when a partial pack is already
> complete and let index-pack validate its contents.

I wonder if we still need this or not. AIUI the original 416 responses came because we were asking for nonsense outside of the range (because the corrupted writes advanced the file too far). The worst case now is that we'd ask for bytes "N-" when the file is only N bytes long, and the server should say "OK, here are your 0 bytes". But maybe there's a server who complains about that.

Show 17 quoted lines
> diff --git a/http.c b/http.c
> index caccf2108e..a0d399b274 100644
> --- a/http.c
> +++ b/http.c
> @@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)
>  	int tmpfile_fd;
>  	int ret = 0;
>  
> +	/* Another downloader may unlink the staging path while we index it. */
> +	tmpfile_fd = xdup(fileno(preq->packfile));
>  	fclose(preq->packfile);
>  	preq->packfile = NULL;
> -
> -	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
> +	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
> +		die_errno("unable to seek local file %s for pack",
> +			  preq->tmpfile.buf);

OK, here we are avoiding the race that it gets unlinked by dup-ing the existing descriptor and seeking back to the start. Makes sense. But then...

Show 14 quoted lines
> @@ -2704,13 +2707,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
>  	else
>  		ip.no_stdout = 1;
>  
> -	if (run_command(&ip)) {
> +	if (run_command(&ip))
>  		ret = -1;
> -		goto cleanup;
> -	}
> -
> -cleanup:
> -	close(tmpfile_fd);
>  	unlink(preq->tmpfile.buf);
>  	return ret;

What is going on with this hunk? We don't really need to jump to cleanup here because we get there directly anyway, and there are no other users of the cleanup label. So that part doesn't seem wrong, but rather unrelated.

More importantly, why don't we need to close tmpfile_fd anymore? We hand it off to run_command(), which will always close it. So I _think_ it was always wrong to close it ourselves here. If so, then could this hunk become a preparatory commit on its own?

This commit is already confusing enough that the more extraneous stuff we can take out of it the better.

Show 18 quoted lines
> @@ -2738,22 +2736,45 @@ struct http_pack_request *new_http_pack_request(
> [...]
> +	/*
> +	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
> +	 * existing file; reopen a newly created file so others may unlink it.
> +	 */
> +	for (;;) {
> +		fd = open(preq->tmpfile.buf, O_RDWR);
> +		if (fd >= 0 || errno != ENOENT)
> +			break;
> +		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
> +		if (fd >= 0) {
> +			close(fd);
> +			continue;
> +		}
> +		if (errno != EEXIST)
> +			break;
> +	}

OK, and this is the opening magic. What's going on with the O_EXCL here, though? We try to open once, and if that fails with ENOENT then we open again. But isn't that racy? Two processes simultaneously try to open(), find the file is not there, and then both try O_EXCL. Only one of them will win, and the other will barf.

I guess that is the reason for the loop, where we will try again over and over until we either pick up somebody else's copy or get our own. And if we get our own, we still close it and try again. And that's the Windows magic described in the commit message.

That is...subtle as hell. I really wonder if it would be worth introducing the basic form of this (just opening once with O_RDWR) and then doing the Windows hackery on top as a separate commit. That would leave the intermediate state subject to racy problems on Windows. But when balancing bisectability versus having a clear human-readable patch, I think I'd rather see it broken up.

Show 5 quoted lines
> +	if (fd < 0) {
> +		error_errno("unable to open local file %s for pack",
> +			    preq->tmpfile.buf);
>  		goto abort;
>  	}

OK, and then we get here if we broke out of the loop due to an error besides ENOENT/EEXIST.

Show 8 quoted lines
> +	prev_posn = lseek(fd, 0, SEEK_END);
> +	if (prev_posn < 0) {
> +		error_errno("unable to seek local file %s for pack",
> +			    preq->tmpfile.buf);
> +		close(fd);
> +		goto abort;
> +	}
> +	preq->packfile = xfdopen(fd, "w");
And then this is the positioning magic to replace O_APPEND. Good.
> diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh

For the record, I don't love that we are using a custom perl script here instead of going through apache (like all of our other tests). But I suspect the apache version would be sufficiently horrific (possibly even worse) that it's not really worth pursuing. Hopefully this perl script (and the accompanying fifo monstrosities) can sit here for eternity un-looked-at by human eyes, just quietly doing their job until the heat death of the universe.

-Peff
Jeff KingJul 26, 2026, 09:21 UTC in reply to Ted Nyman on lore

Re: [PATCH v5 0/3] packfile URIs: support concurrent downloads

On Sat, Jul 25, 2026 at 11:44:45PM -0700, Ted Nyman wrote:
Show 8 quoted lines
> Changes since v4:
> 
>   * Clarify that the first --index-pack-arg specifies the command and
>     subsequent instances specify its arguments.
>   * Drop assumptions about which concurrent response reaches the
>     staging file first. Either write order exercises the same
>     overlapping-download behavior.
>   * No production code changes.

Thanks. I hadn't really reviewed the code in v4 carefully, but I did so for v5. I _think_ it is all correct, but it there are a few confusing bits in the middle patch that might be worth breaking apart for readability.

I could live with it as-is, though.
-Peff
Ted NymanJul 26, 2026, 10:04 UTC in reply to Jeff King on lore

Re: [PATCH v5 2/3] http: avoid concurrent appends to partial packs

On Sun, Jul 26, 2026 at 05:20:27AM -0400, Jeff King wrote:
> I wonder if we still need this or not.

I think so, but wouldn't bet the farm on it. A concurrent downloader can complete the staging file before another downloader issues its Range request. That request then starts exactly at EOF, so the server can respond with 416. The existing regression test exercises that case, and we still need to let index-pack validate the completed local pack.

> More importantly, why don't we need to close tmpfile_fd anymore? We hand
> it off to run_command(), which will always close it. So I _think_ it was
> always wrong to close it ourselves here. If so, then could this hunk
> become a preparatory commit on its own?

You're right: run_command() already closes ip.in, so the old close(tmpfile_fd) was a double-close. That cleanup is independent, and I can pull it into a preparatory patch if that would make the series easier to follow.

> That is...subtle as hell. I really wonder if it would be worth
> introducing the basic form of this (just opening once with O_RDWR) and
> then doing the Windows hackery on top as a separate commit.

I'm certainly not an expert on the Windows side, so I had to track this down in the MinGW open() wrapper. The existing-file O_RDWR path includes FILE_SHARE_DELETE, while creating a new file falls back to _wopen() without it. The loop creates the file with O_EXCL if needed, closes that descriptor, and retries through the existing-file path; a racing creator that sees EEXIST also retries.

I kept those pieces together to avoid an intermediate state without the required sharing behavior on MinGW, but I'm happy to split them if you think it would be clearer.

> Hopefully this perl script (and the accompanying fifo monstrosities)
> can sit here for eternity un-looked-at by human eyes, just quietly
> doing their job until the heat death of the universe.
I thought you, of all people, might appreciate a little more Perl. ;-)

Thanks, Ted

Jeff KingJul 26, 2026, 10:27 UTC in reply to Ted Nyman on lore

Re: [PATCH v5 2/3] http: avoid concurrent appends to partial packs

On Sun, Jul 26, 2026 at 03:04:21AM -0700, Ted Nyman wrote:
Show 8 quoted lines
> On Sun, Jul 26, 2026 at 05:20:27AM -0400, Jeff King wrote:
> > I wonder if we still need this or not.
> 
> I think so, but wouldn't bet the farm on it. A concurrent downloader can
> complete the staging file before another downloader issues its Range
> request. That request then starts exactly at EOF, so the server can
> respond with 416. The existing regression test exercises that case, and
> we still need to let index-pack validate the completed local pack.

Yeah, the big question there for me is whether a server would return a 416 in such a case. It seems reasonable that a client might know it has N bytes but not the full size, and ask for "N-", expecting to get some equivalent of a 0-byte read(). Whereas a 416 does not make it clear at all whether the range is nonsense, or if you happened to be at EOF.

But sadly we do not seem to live in that world, based on a few tests. So I agree we do need it to cover that edge case.

If you are splitting things out of the patch, can we do the same for this 416 handling? It is already a problem even without concurrency if you happen to get the full file but then fail for other reasons before indexing the pack (transient system errors, etc).

To be clear, I can live with things as they are and I don't want to make too much work for you in splitting. But I'm hoping that feeding it to an electronic friend could do that split without much effort.

Show 10 quoted lines
> > That is...subtle as hell. I really wonder if it would be worth
> > introducing the basic form of this (just opening once with O_RDWR) and
> > then doing the Windows hackery on top as a separate commit.
> 
> I'm certainly not an expert on the Windows side, so I had to track this
> down in the MinGW open() wrapper. The existing-file O_RDWR path includes
> FILE_SHARE_DELETE, while creating a new file falls back to _wopen()
> without it. The loop creates the file with O_EXCL if needed, closes that
> descriptor, and retries through the existing-file path; a racing creator
> that sees EEXIST also retries.

Yeah, I understand it now after reading the commit message and the code several times. The loop is what I think is subtle, but I can't see a more obvious way of writing it that deals with all of the possible combinations and races.

> I kept those pieces together to avoid an intermediate state without the
> required sharing behavior on MinGW, but I'm happy to split them if you
> think it would be clearer.

IMHO it is the lesser of two evils. Not because we won't eventually end up with the subtle loop, but because it helps make the desired change in the "simpler" version much easier to see. IOW, the Windows patch is always going to be confusing, but we can at least salvage the original.

Show 5 quoted lines
> > Hopefully this perl script (and the accompanying fifo monstrosities)
> > can sit here for eternity un-looked-at by human eyes, just quietly
> > doing their job until the heat death of the universe.
> 
> I thought you, of all people, might appreciate a little more Perl. ;-)

Well, if we have to write something like that, obviously Perl is the right choice. ;) It's more the custom socket handling. E.g., there are sometimes subtle issues around listen-port allocation, especially when we run the same test concurrently with --stress. Your script solves it by asking for a dynamic port and then passing that back to the caller over the fifo. That should work reliably, I think, it's just not our usual solution (but usually we are constrained to having to tell apache the correct port up front, which means we need to pick an unambiguous one ourselves).

-Peff
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 0/6] packfile URIs: support concurrent downloads

Packfile URI and dumb HTTP downloads stage packs at objects/pack/pack-<hash>.pack.temp so an interrupted transfer can resume. Opening that file in append mode forces every write to its current end. Two Git processes fetching the same pack into one object database can therefore append duplicate data and corrupt the pack.

The first patch separates the unrelated --index-pack-arg documentation and error-message correction requested during review.

The second patch fixes an existing double-close when finish_http_pack_request() passes its staging-file descriptor to index-pack. start_command() already takes ownership of that descriptor, including when starting the child fails.

The third patch handles a completed partial pack independently of concurrent downloads. A previous attempt can finish the transfer but fail before indexing it; retrying then requests a range starting at EOF. Servers may respond with HTTP 416 in that case. Treat the response as a completed download and let index-pack validate the pack.

The fourth patch keeps the predictable staging name but removes append mode. Each downloader seeks once to the current end, requests the corresponding Range, and writes using its own descriptor offset. Since the staging key must identify immutable pack contents, overlapping responses write identical bytes at identical offsets. There is no need for pwrite(2) or cross-process coordination, and resumption continues to work for both packfile URI and ordinary dumb HTTP downloads.

The fifth patch handles the additional MinGW sharing requirement. Its non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing staging file exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the path.

The final patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process. Read only the prefix and hash so any following fsck output remains available to fetch-pack.

The tests cover resumption, a completed partial returning 416, overlapping downloads, unlinking the staging path while index-pack holds its descriptor, and a pre-existing .keep file. The completed-partial and unlink tests do not require FIFOs, so they can run on MinGW even though the concurrent-download test is skipped there.

Changes since v5:
* Split the existing double-close fix, HTTP 416 handling, generic
  concurrent-download fix, and Windows sharing fix into separate
  patches.
* Replace the FIFO-based concurrent HTTP 416 test with a standalone
  completed-partial test. Besides simplifying the test, this covers the
  non-concurrent interrupted-download case directly.
* Keep the final production code unchanged.

Each patch passes t5550-http-fetch-dumb.sh. The final series also passes t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs with 12 parallel stress jobs.

The v5 discussion is at:
https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/
Ted Nyman (6):
  http-fetch: correct --index-pack-arg documentation
  http: avoid closing index-pack input twice
  http: accept HTTP 416 for complete partial packs
  http: avoid concurrent appends to partial packs
  http: permit unlinking partial packs on Windows
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  14 +-
 fetch-pack.c                      |  33 ++---
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 +++++---
 t/t5550-http-fetch-dumb.sh        | 204 ++++++++++++++++++++++++++++++
 t/t5702-protocol-v2.sh            |  31 +++++
 8 files changed, 305 insertions(+), 46 deletions(-)
Range-diff against v5:
1:  a79af009ea = 1:  b5050a88ca http-fetch: correct --index-pack-arg documentation
-:  ---------- > 2:  28662b0fd8 http: avoid closing index-pack input twice
-:  ---------- > 3:  677e5399eb http: accept HTTP 416 for complete partial packs
2:  d9667c93b0 ! 4:  7a83eb7091 http: avoid concurrent appends to partial packs
    @@ Commit message
         pack contents, overlapping responses write the same bytes at the same
         offsets instead of extending the file with duplicate data.
     
    -    MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
    -    existing file. Create a missing partial pack exclusively, close it, and
    -    reopen it without O_CREAT so every retained descriptor permits another
    -    downloader to unlink the staging path. Duplicate that descriptor for
    -    index-pack instead of reopening the path after closing the stream;
    -    index-pack installs its own pack and the shared staging file is only
    -    unlinked, never renamed. Accept HTTP 416 when a partial pack is already
    -    complete and let index-pack validate its contents.
    +    Duplicate the staging descriptor for index-pack instead of reopening the
    +    path after closing the stream. Another downloader may unlink the staging
    +    path before indexing begins, but index-pack can still read the retained
    +    descriptor.
     
    -    Exercise resumed transfers, EOF ranges, overlapping 200 and 206
    -    responses, and unlinking the staging path while index-pack still holds
    -    its descriptor. Clarify the staging-key documentation.
    +    Exercise resumed transfers and overlapping 200 and 206 responses, and
    +    clarify the staging-key documentation.
     
         Signed-off-by: Ted Nyman <tnyman@openai.com>
     
    @@ Documentation/git-http-fetch.adoc: commit-id::
      
      --index-pack-arg=<arg>::
     
    - ## http-fetch.c ##
    -@@ http-fetch.c: static void fetch_single_packfile(struct object_id *packfile_hash,
    - 
    - 	if (start_active_slot(preq->slot)) {
    - 		run_active_slot(preq->slot);
    --		if (results.curl_result != CURLE_OK) {
    -+		if (results.curl_result != CURLE_OK &&
    -+		    results.http_code != 416) {
    - 			struct url_info url;
    - 			char *nurl = url_normalize(preq->url, &url);
    - 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
    -
    - ## http-push.c ##
    -@@ http-push.c: static void finish_request(struct transfer_request *request)
    - 
    - 	} else if (request->state == RUN_FETCH_PACKED) {
    - 		int fail = 1;
    --		if (request->curl_result != CURLE_OK) {
    -+		if (request->curl_result != CURLE_OK &&
    -+		    request->http_code != 416) {
    - 			fprintf(stderr, "Unable to get pack file %s\n%s",
    - 				request->url, curl_errorstr);
    - 		} else {
    -
    - ## http-walker.c ##
    -@@ http-walker.c: static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
    - 
    - 	if (start_active_slot(preq->slot)) {
    - 		run_active_slot(preq->slot);
    --		if (results.curl_result != CURLE_OK) {
    -+		if (results.curl_result != CURLE_OK &&
    -+		    results.http_code != 416) {
    - 			error("Unable to get pack file %s\n%s", preq->url,
    - 			      curl_errorstr);
    - 			goto abort;
    -
      ## http.c ##
     @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)
      	int tmpfile_fd;
    @@ http.c: int finish_http_pack_request(struct http_pack_request *preq)
      
      	ip.git_cmd = 1;
      	ip.in = tmpfile_fd;
    -@@ http.c: int finish_http_pack_request(struct http_pack_request *preq)
    - 	else
    - 		ip.no_stdout = 1;
    - 
    --	if (run_command(&ip)) {
    -+	if (run_command(&ip))
    - 		ret = -1;
    --		goto cleanup;
    --	}
    --
    --cleanup:
    --	close(tmpfile_fd);
    - 	unlink(preq->tmpfile.buf);
    - 	return ret;
    - }
     @@ http.c: struct http_pack_request *new_http_pack_request(
      struct http_pack_request *new_direct_http_pack_request(
      	const unsigned char *packed_git_hash, char *url)
    @@ http.c: struct http_pack_request *new_http_pack_request(
     -	if (!preq->packfile) {
     -		error("Unable to open local file %s for pack",
     -		      preq->tmpfile.buf);
    -+	/*
    -+	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
    -+	 * existing file; reopen a newly created file so others may unlink it.
    -+	 */
    -+	for (;;) {
    -+		fd = open(preq->tmpfile.buf, O_RDWR);
    -+		if (fd >= 0 || errno != ENOENT)
    -+			break;
    -+		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
    -+		if (fd >= 0) {
    -+			close(fd);
    -+			continue;
    -+		}
    -+		if (errno != EEXIST)
    -+			break;
    -+	}
    ++	fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT, 0666);
     +	if (fd < 0) {
     +		error_errno("unable to open local file %s for pack",
     +			    preq->tmpfile.buf);
    - 		goto abort;
    - 	}
    ++		goto abort;
    ++	}
     +	prev_posn = lseek(fd, 0, SEEK_END);
     +	if (prev_posn < 0) {
     +		error_errno("unable to seek local file %s for pack",
     +			    preq->tmpfile.buf);
     +		close(fd);
    -+		goto abort;
    -+	}
    + 		goto abort;
    + 	}
     +	preq->packfile = xfdopen(fd, "w");
      
      	preq->slot = get_active_slot();
    @@ http.c: struct http_pack_request *new_direct_http_pack_request(
      				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
     
      ## t/t5550-http-fetch-dumb.sh ##
    -@@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
    - 	git -C packfileclient cat-file -e "$HASH"
    +@@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile accepts an already complete partial'
    + 	git -C packfileclient-complete cat-file -e "$HASH"
      '
      
     +test_expect_success 'http-fetch --packfile resumes a partial download' '
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	git -C packfileclient-resume cat-file -e "$HASH"
     +'
     +
    -+test_expect_success 'http-fetch --packfile permits unlink while indexing' '
    -+	git init packfileclient-unlink &&
    -+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
    -+		ls objects/pack/pack-*.pack) &&
    -+	tmpfile="packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
    -+	write_script git-unlink-index-pack <<-\EOF &&
    -+	test -f "$GIT_TEST_PACK_TEMP" || exit 1
    -+	rm "$GIT_TEST_PACK_TEMP" || exit 1
    -+	exec git index-pack "$@"
    -+	EOF
    -+	test_when_finished "rm -f git-unlink-index-pack" &&
    -+	PATH="$TRASH_DIRECTORY:$PATH" \
    -+	GIT_TEST_PACK_TEMP="$TRASH_DIRECTORY/$tmpfile" \
    -+	git -C packfileclient-unlink http-fetch --packfile="$ARBITRARY" \
    -+		--index-pack-arg=unlink-index-pack \
    -+		--index-pack-arg=--stdin --index-pack-arg=--keep \
    -+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
    -+	test_path_is_missing "$tmpfile" &&
    -+	git -C packfileclient-unlink cat-file -e "$HASH"
    -+'
    -+
    -+test_expect_success PIPE 'concurrent http-fetch --packfile accepts a complete partial' '
    -+	git init packfileclient-concurrent &&
    -+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
    -+		ls objects/pack/pack-*.pack) &&
    -+	packhash=$(basename "$p" .pack) &&
    -+	packhash=${packhash#pack-} &&
    -+	tmpfile="packfileclient-concurrent/.git/objects/pack/pack-$packhash.pack.temp" &&
    -+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
    -+	mkfifo first-ready first-continue &&
    -+	exec 8<>first-ready &&
    -+	exec 9<>first-continue &&
    -+	write_script git-wait-index-pack <<-\EOF &&
    -+	echo ready >"$GIT_TEST_WAIT_READY" &&
    -+	read continue <"$GIT_TEST_WAIT_CONTINUE" &&
    -+	exec git index-pack "$@"
    -+	EOF
    -+	{
    -+		(
    -+			if ! PATH="$TRASH_DIRECTORY:$PATH" \
    -+			GIT_TEST_WAIT_READY="$TRASH_DIRECTORY/first-ready" \
    -+			GIT_TEST_WAIT_CONTINUE="$TRASH_DIRECTORY/first-continue" \
    -+			GIT_TRACE_CURL="$TRASH_DIRECTORY/first.trace" \
    -+			git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
    -+				--index-pack-arg=wait-index-pack \
    -+				--index-pack-arg=--stdin --index-pack-arg=--keep \
    -+				"$HTTPD_URL/dumb/repo_pack.git/$p" >first.out
    -+			then
    -+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
    -+				exit 1
    -+			fi
    -+		) &
    -+		first_pid=$!
    -+	} &&
    -+	test_when_finished "
    -+		echo continue >&9
    -+		kill $first_pid 2>/dev/null || :
    -+		wait $first_pid 2>/dev/null || :
    -+		exec 8>&-
    -+		exec 9>&-
    -+		rm -f first-ready first-continue git-wait-index-pack
    -+	" &&
    -+	read ready <&8 &&
    -+	test "$ready" = ready &&
    -+	GIT_TRACE_CURL="$TRASH_DIRECTORY/second.trace" \
    -+	git -C packfileclient-concurrent http-fetch --packfile="$packhash" \
    -+		--index-pack-arg=index-pack \
    -+		--index-pack-arg=--stdin --index-pack-arg=--keep \
    -+		"$HTTPD_URL/dumb/repo_pack.git/$p" >second.out &&
    -+	echo continue >&9 &&
    -+	wait "$first_pid" &&
    -+	printf "pack\t%s\n" "$packhash" >expect &&
    -+	test_cmp expect first.out &&
    -+	printf "keep\t%s\n" "$packhash" >expect &&
    -+	test_cmp expect second.out &&
    -+	test_grep "Range: bytes=64-" first.trace &&
    -+	test_grep "Range: bytes=[0-9]*-" second.trace &&
    -+	test_grep "416 Requested Range Not Satisfiable" second.trace &&
    -+	test_path_is_missing "$tmpfile" &&
    -+	git -C packfileclient-concurrent cat-file -e "$HASH"
    -+'
    -+
     +test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
     +	git init packfileclient-overlap &&
     +	blob=$(test-tool genrandom pack-overlap 2m |
-:  ---------- > 5:  87a20ac80f http: permit unlinking partial packs on Windows
3:  fee6f292cb = 6:  be9e2fe273 fetch-pack: accept "pack" output for packfile URIs
base-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 1/6] http-fetch: correct --index-pack-arg documentation

The --packfile mode accepts one --index-pack-arg=<arg> option per argument passed to index-pack, but its documentation and option dependency errors still refer to the plural --index-pack-args form.

Correct the spelling and describe the repeatable per-argument form.
Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc | 9 +++++----
 http-fetch.c                      | 4 ++--
 2 files changed, 7 insertions(+), 6 deletions(-)
Show changes to 2 files +7 −5

Documentation/git-http-fetch.adoc, http-fetch.c

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 2200f073c4..12036e65e9 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -50,11 +50,12 @@ commit-id::
 	URL and uses index-pack to generate corresponding .idx and .keep files.
 	The hash is used to determine the name of the temporary file and is
 	arbitrary. The output of index-pack is printed to stdout. Requires
-	--index-pack-args.
+	one or more --index-pack-arg options.
 
---index-pack-args=<args>::
-	For internal use only. The command to run on the contents of the
-	downloaded pack. Arguments are URL-encoded separated by spaces.
+--index-pack-arg=<arg>::
+	For internal use only. The first instance specifies the command run on
+	the contents of the downloaded pack. Subsequent instances specify its
+	arguments.
 
 --recover::
 	Verify that everything reachable from target is fetched.  Used after
diff --git a/http-fetch.c b/http-fetch.c
index f9b6ecb061..601a77c3c1 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -155,7 +155,7 @@ int cmd_main(int argc, const char **argv)
 
 	if (packfile) {
 		if (!index_pack_args.nr)
-			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
+			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
 
 		fetch_single_packfile(&packfile_hash, argv[arg],
 				      index_pack_args.v);
@@ -164,7 +164,7 @@ int cmd_main(int argc, const char **argv)
 	}
 
 	if (index_pack_args.nr)
-		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
+		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
 
 	if (commits_on_stdin) {
 		commits = walker_targets_stdin(&commit_id, &write_ref);
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 2/6] http: avoid closing index-pack input twice

finish_http_pack_request() passes its staging-file descriptor to index-pack through child_process.in. start_command() takes ownership of a supplied descriptor and closes it, even when starting the child fails.

Do not close the descriptor again after run_command() returns.
Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 http.c | 7 +------
 1 file changed, 1 insertion(+), 6 deletions(-)
Show changes to http.c +1 −6
diff --git a/http.c b/http.c
index caccf2108e..89a1ccc6d2 100644
--- a/http.c
+++ b/http.c
@@ -2704,13 +2704,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	else
 		ip.no_stdout = 1;
 
-	if (run_command(&ip)) {
+	if (run_command(&ip))
 		ret = -1;
-		goto cleanup;
-	}
-
-cleanup:
-	close(tmpfile_fd);
 	unlink(preq->tmpfile.buf);
 	return ret;
 }
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 3/6] http: accept HTTP 416 for complete partial packs

A resumed pack request may already have all bytes of the remote pack. A server can respond to the resulting Range request with HTTP 416 instead of returning an empty response.

Accept that response in each pack-download caller and let index-pack validate the completed staging file. This can happen without concurrent downloads when a previous attempt completed the transfer but failed before indexing it.

Add a regression test that seeds a complete partial pack and checks that http-fetch indexes it after the server returns HTTP 416.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 http-fetch.c               |  3 ++-
 http-push.c                |  3 ++-
 http-walker.c              |  3 ++-
 t/t5550-http-fetch-dumb.sh | 19 +++++++++++++++++++
 4 files changed, 25 insertions(+), 3 deletions(-)
Show changes to 4 files +25 −3

http-fetch.c, http-push.c, http-walker.c, t/t5550-http-fetch-dumb.sh

diff --git a/http-fetch.c b/http-fetch.c
index 601a77c3c1..05f68f306a 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			struct url_info url;
 			char *nurl = url_normalize(preq->url, &url);
 			if (!nurl || !git_env_bool("GIT_TRACE_REDACT", 1)) {
diff --git a/http-push.c b/http-push.c
index 60f6f8f054..ef8abe3908 100644
--- a/http-push.c
+++ b/http-push.c
@@ -595,7 +595,8 @@ static void finish_request(struct transfer_request *request)
 
 	} else if (request->state == RUN_FETCH_PACKED) {
 		int fail = 1;
-		if (request->curl_result != CURLE_OK) {
+		if (request->curl_result != CURLE_OK &&
+		    request->http_code != 416) {
 			fprintf(stderr, "Unable to get pack file %s\n%s",
 				request->url, curl_errorstr);
 		} else {
diff --git a/http-walker.c b/http-walker.c
index b58a3b2a92..abafca84d6 100644
--- a/http-walker.c
+++ b/http-walker.c
@@ -451,7 +451,8 @@ static int http_fetch_pack(struct walker *walker, struct alt_base *repo,
 
 	if (start_active_slot(preq->slot)) {
 		run_active_slot(preq->slot);
-		if (results.curl_result != CURLE_OK) {
+		if (results.curl_result != CURLE_OK &&
+		    results.http_code != 416) {
 			error("Unable to get pack file %s\n%s", preq->url,
 			      curl_errorstr);
 			goto abort;
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index f00eeae48f..698bbb3160 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -293,6 +293,25 @@ test_expect_success 'http-fetch --packfile' '
 	git -C packfileclient cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile accepts an already complete partial' '
+	git init packfileclient-complete &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	packhash=$(basename "$p" .pack) &&
+	packhash=${packhash#pack-} &&
+	tmpfile="packfileclient-complete/.git/objects/pack/pack-$packhash.pack.temp" &&
+	cp "$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" "$tmpfile" &&
+	chmod u+w "$tmpfile" &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/complete.trace" \
+	git -C packfileclient-complete http-fetch --packfile="$packhash" \
+		--index-pack-arg=index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_grep "416 Requested Range Not Satisfiable" complete.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-complete cat-file -e "$HASH"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 4/6] http: avoid concurrent appends to partial packs

Pack requests stage downloads in a predictable partial-pack file so an interrupted transfer can be resumed. Both packfile URI and ordinary dumb HTTP requests use this staging path. Opening it in append mode forces each write to the current end of the file, so concurrent responses can append duplicate data and corrupt the pack.

Open the partial pack read-write without O_APPEND and seek once to its current end. Each downloader then retains the offset matching the Range it requested. Because the staging key must uniquely identify immutable pack contents, overlapping responses write the same bytes at the same offsets instead of extending the file with duplicate data.

Duplicate the staging descriptor for index-pack instead of reopening the path after closing the stream. Another downloader may unlink the staging path before indexing begins, but index-pack can still read the retained descriptor.

Exercise resumed transfers and overlapping 200 and 206 responses, and clarify the staging-key documentation.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 Documentation/git-http-fetch.adoc |   5 +-
 http.c                            |  34 ++++---
 t/t5550-http-fetch-dumb.sh        | 164 ++++++++++++++++++++++++++++++
 3 files changed, 187 insertions(+), 16 deletions(-)
Show changes to 3 files +187 −16

Documentation/git-http-fetch.adoc, http.c, t/t5550-http-fetch-dumb.sh

diff --git a/Documentation/git-http-fetch.adoc b/Documentation/git-http-fetch.adoc
index 12036e65e9..45e0d3d07c 100644
--- a/Documentation/git-http-fetch.adoc
+++ b/Documentation/git-http-fetch.adoc
@@ -48,8 +48,9 @@ commit-id::
 	line (which is not expected in
 	this case), 'git http-fetch' fetches the packfile directly at the given
 	URL and uses index-pack to generate corresponding .idx and .keep files.
-	The hash is used to determine the name of the temporary file and is
-	arbitrary. The output of index-pack is printed to stdout. Requires
+	The hash is used to determine the name of the temporary file. It need
+	not be the pack hash, but it must uniquely identify the pack contents
+	for resumption. The output of index-pack is printed to stdout. Requires
 	one or more --index-pack-arg options.
 
 --index-pack-arg=<arg>::
diff --git a/http.c b/http.c
index 89a1ccc6d2..ad07ef3549 100644
--- a/http.c
+++ b/http.c
@@ -2688,10 +2688,13 @@ int finish_http_pack_request(struct http_pack_request *preq)
 	int tmpfile_fd;
 	int ret = 0;
 
+	/* Another downloader may unlink the staging path while we index it. */
+	tmpfile_fd = xdup(fileno(preq->packfile));
 	fclose(preq->packfile);
 	preq->packfile = NULL;
-
-	tmpfile_fd = xopen(preq->tmpfile.buf, O_RDONLY);
+	if (lseek(tmpfile_fd, 0, SEEK_SET) < 0)
+		die_errno("unable to seek local file %s for pack",
+			  preq->tmpfile.buf);
 
 	ip.git_cmd = 1;
 	ip.in = tmpfile_fd;
@@ -2733,22 +2736,30 @@ struct http_pack_request *new_http_pack_request(
 struct http_pack_request *new_direct_http_pack_request(
 	const unsigned char *packed_git_hash, char *url)
 {
-	off_t prev_posn = 0;
+	off_t prev_posn;
 	struct http_pack_request *preq;
+	int fd;
 
 	CALLOC_ARRAY(preq, 1);
 	strbuf_init(&preq->tmpfile, 0);
-
 	preq->url = url;
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	preq->packfile = fopen(preq->tmpfile.buf, "a");
-	if (!preq->packfile) {
-		error("Unable to open local file %s for pack",
-		      preq->tmpfile.buf);
+	fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT, 0666);
+	if (fd < 0) {
+		error_errno("unable to open local file %s for pack",
+			    preq->tmpfile.buf);
+		goto abort;
+	}
+	prev_posn = lseek(fd, 0, SEEK_END);
+	if (prev_posn < 0) {
+		error_errno("unable to seek local file %s for pack",
+			    preq->tmpfile.buf);
+		close(fd);
 		goto abort;
 	}
+	preq->packfile = xfdopen(fd, "w");
 
 	preq->slot = get_active_slot();
 	preq->headers = object_request_headers();
@@ -2757,12 +2768,7 @@ struct http_pack_request *new_direct_http_pack_request(
 	curl_easy_setopt(preq->slot->curl, CURLOPT_URL, preq->url);
 	curl_easy_setopt(preq->slot->curl, CURLOPT_HTTPHEADER, preq->headers);
 
-	/*
-	 * If there is data present from a previous transfer attempt,
-	 * resume where it left off
-	 */
-	prev_posn = ftello(preq->packfile);
-	if (prev_posn>0) {
+	if (prev_posn > 0) {
 		if (http_is_verbose)
 			fprintf(stderr,
 				"Resuming fetch of pack %s at byte %"PRIuMAX"\n",
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index 698bbb3160..86b9d87ef5 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -312,6 +312,170 @@ test_expect_success 'http-fetch --packfile accepts an already complete partial'
 	git -C packfileclient-complete cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile resumes a partial download' '
+	git init packfileclient-resume &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-resume/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	test_copy_bytes 64 <"$HTTPD_DOCUMENT_ROOT_PATH/repo_pack.git/$p" >"$tmpfile" &&
+	GIT_TRACE_CURL="$TRASH_DIRECTORY/resume.trace" \
+	git -C packfileclient-resume http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=index-pack --index-pack-arg=--stdin \
+		--index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_grep "Range: bytes=64-" resume.trace &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-resume cat-file -e "$HASH"
+'
+
+test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
+	git init packfileclient-overlap &&
+	blob=$(test-tool genrandom pack-overlap 2m |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			hash-object -w --stdin) &&
+	packhash=$(printf "%s\n" "$blob" |
+		git -C "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git \
+			pack-objects "$TRASH_DIRECTORY/overlap-pack") &&
+	pack="$TRASH_DIRECTORY/overlap-pack-$packhash.pack" &&
+	tmpfile="packfileclient-overlap/.git/objects/pack/pack-$packhash.pack.temp" &&
+	mkfifo server-ready first-ready &&
+	exec 7<>server-ready &&
+	exec 8<>first-ready &&
+	write_script slow-pack-server "$PERL_PATH" <<-\EOF &&
+	use strict;
+	use warnings;
+	use IO::Socket::INET;
+
+	my ($packfile, $server_ready, $first_ready) = @ARGV;
+	my $completed = 0;
+	END {
+		if (!$completed) {
+			signal_ready($server_ready, "failed");
+			signal_ready($first_ready, "failed");
+		}
+	}
+
+	$SIG{ALRM} = sub { die "timed out serving concurrent pack requests\n" };
+	alarm 60;
+
+	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
+	my $pack = do { local $/; <$in> };
+	close($in) or die "close $packfile: $!";
+	my $server = IO::Socket::INET->new(LocalAddr => "127.0.0.1",
+		LocalPort => 0, Proto => "tcp", Listen => 2, ReuseAddr => 1)
+		or die "listen: $!";
+
+	sub signal_ready {
+		my ($file, $value) = @_;
+		open(my $out, ">", $file) or die "open $file: $!";
+		print $out "$value\n" or die "write $file: $!";
+		close($out) or die "close $file: $!";
+	}
+
+	sub write_all {
+		my ($out, $data) = @_;
+		my $offset = 0;
+		while ($offset < length($data)) {
+			my $written = syswrite($out, $data,
+				length($data) - $offset, $offset);
+			defined($written) && $written or die "write response: $!";
+			$offset += $written;
+		}
+	}
+
+	sub start_response {
+		my $out = $server->accept() or die "accept: $!";
+		<$out> or die "read request: $!";
+		my $start = 0;
+		while (<$out>) {
+			last if /^\r?\n$/;
+			$start = $1 if /^Range: bytes=(\d+)-/i;
+		}
+		$start < length($pack) or die "invalid range $start";
+		my $length = length($pack) - $start;
+		my $middle = int($length / 2);
+		my $status = $start ? "206 Partial Content" : "200 OK";
+		my $headers = "HTTP/1.1 $status\r\n" .
+			"Content-Length: $length\r\n" .
+			($start ? "Content-Range: bytes $start-" .
+				(length($pack) - 1) . "/" . length($pack) . "\r\n" : "") .
+			"Connection: close\r\n\r\n";
+		write_all($out, $headers);
+		write_all($out, substr($pack, $start, $middle));
+		return ($out, $start + $middle);
+	}
+
+	signal_ready($server_ready, $server->sockport());
+	my ($first, $first_pos) = start_response();
+	signal_ready($first_ready, "ready");
+	my ($second, $second_pos) = start_response();
+	write_all($first, substr($pack, $first_pos));
+	write_all($second, substr($pack, $second_pos));
+	close($first) or die "close first response: $!";
+	close($second) or die "close second response: $!";
+	$completed = 1;
+	alarm 0;
+	EOF
+	{
+		"$TRASH_DIRECTORY/slow-pack-server" "$pack" \
+			"$TRASH_DIRECTORY/server-ready" \
+			"$TRASH_DIRECTORY/first-ready" >server.log 2>&1 &
+		server_pid=$!
+	} &&
+	test_when_finished "
+		kill $server_pid 2>/dev/null || :
+		wait $server_pid 2>/dev/null || :
+		exec 7>&-
+		exec 8>&-
+		rm -f server-ready first-ready slow-pack-server
+	" &&
+	read port <&7 &&
+	url="http://127.0.0.1:$port/pack" &&
+	{
+		(
+			if ! GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-first.trace" \
+			GIT_TRACE_CURL_NO_DATA=1 \
+			git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+				--index-pack-arg=index-pack \
+				--index-pack-arg=--stdin --index-pack-arg=--keep \
+				"$url" >first.out
+			then
+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
+				exit 1
+			fi
+		) &
+		first_pid=$!
+	} &&
+	test_when_finished "
+		kill $first_pid 2>/dev/null || :
+		wait $first_pid 2>/dev/null || :
+	" &&
+	read ready <&8 &&
+	test "$ready" = ready &&
+	test_path_is_file "$tmpfile" &&
+	{
+		GIT_TRACE_CURL="$TRASH_DIRECTORY/overlap-second.trace" \
+		GIT_TRACE_CURL_NO_DATA=1 \
+		git -C packfileclient-overlap http-fetch --packfile="$packhash" \
+			--index-pack-arg=index-pack \
+			--index-pack-arg=--stdin --index-pack-arg=--keep \
+			"$url" >second.out &
+		second_pid=$!
+	} &&
+	test_when_finished "
+		kill $second_pid 2>/dev/null || :
+		wait $second_pid 2>/dev/null || :
+	" &&
+	wait "$second_pid" &&
+	wait "$first_pid" &&
+	wait "$server_pid" &&
+	printf "keep\t%s\npack\t%s\n" "$packhash" "$packhash" | sort >expect &&
+	sort first.out second.out >actual &&
+	test_cmp expect actual &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-overlap cat-file -e "$blob"
+'
+
 test_expect_success 'fetch notices corrupt pack' '
 	cp -R "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
 	(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_bad1.git &&
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 5/6] http: permit unlinking partial packs on Windows

On Windows, an open file must permit FILE_SHARE_DELETE before another process can unlink it. MinGW's non-append O_RDWR open enables that sharing mode only for an existing file; adding O_CREAT falls back to _wopen(), which cannot set it.

First try opening the partial pack without O_CREAT. If it does not exist, create it exclusively, close that descriptor, and retry through the existing-file path. A racing creator retries after EEXIST.

This ensures that every retained descriptor permits another downloader to unlink the staging path. Add an unlink-while-indexing test that does not require FIFOs and can therefore run on MinGW.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 http.c                     | 17 ++++++++++++++++-
 t/t5550-http-fetch-dumb.sh | 21 +++++++++++++++++++++
 2 files changed, 37 insertions(+), 1 deletion(-)
Show changes to 2 files +37 −1

http.c, t/t5550-http-fetch-dumb.sh

diff --git a/http.c b/http.c
index ad07ef3549..a0d399b274 100644
--- a/http.c
+++ b/http.c
@@ -2746,7 +2746,22 @@ struct http_pack_request *new_direct_http_pack_request(
 
 	odb_pack_name(the_repository, &preq->tmpfile, packed_git_hash, "pack");
 	strbuf_addstr(&preq->tmpfile, ".temp");
-	fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT, 0666);
+	/*
+	 * MinGW's non-append O_RDWR open grants FILE_SHARE_DELETE only for an
+	 * existing file; reopen a newly created file so others may unlink it.
+	 */
+	for (;;) {
+		fd = open(preq->tmpfile.buf, O_RDWR);
+		if (fd >= 0 || errno != ENOENT)
+			break;
+		fd = open(preq->tmpfile.buf, O_RDWR | O_CREAT | O_EXCL, 0666);
+		if (fd >= 0) {
+			close(fd);
+			continue;
+		}
+		if (errno != EEXIST)
+			break;
+	}
 	if (fd < 0) {
 		error_errno("unable to open local file %s for pack",
 			    preq->tmpfile.buf);
diff --git a/t/t5550-http-fetch-dumb.sh b/t/t5550-http-fetch-dumb.sh
index 86b9d87ef5..b5758f1c9c 100755
--- a/t/t5550-http-fetch-dumb.sh
+++ b/t/t5550-http-fetch-dumb.sh
@@ -328,6 +328,27 @@ test_expect_success 'http-fetch --packfile resumes a partial download' '
 	git -C packfileclient-resume cat-file -e "$HASH"
 '
 
+test_expect_success 'http-fetch --packfile permits unlink while indexing' '
+	git init packfileclient-unlink &&
+	p=$(cd "$HTTPD_DOCUMENT_ROOT_PATH"/repo_pack.git &&
+		ls objects/pack/pack-*.pack) &&
+	tmpfile="packfileclient-unlink/.git/objects/pack/pack-$ARBITRARY.pack.temp" &&
+	write_script git-unlink-index-pack <<-\EOF &&
+	test -f "$GIT_TEST_PACK_TEMP" || exit 1
+	rm "$GIT_TEST_PACK_TEMP" || exit 1
+	exec git index-pack "$@"
+	EOF
+	test_when_finished "rm -f git-unlink-index-pack" &&
+	PATH="$TRASH_DIRECTORY:$PATH" \
+	GIT_TEST_PACK_TEMP="$TRASH_DIRECTORY/$tmpfile" \
+	git -C packfileclient-unlink http-fetch --packfile="$ARBITRARY" \
+		--index-pack-arg=unlink-index-pack \
+		--index-pack-arg=--stdin --index-pack-arg=--keep \
+		"$HTTPD_URL/dumb/repo_pack.git/$p" >out &&
+	test_path_is_missing "$tmpfile" &&
+	git -C packfileclient-unlink cat-file -e "$HASH"
+'
+
 test_expect_success PERL,PIPE 'concurrent http-fetch --packfile cannot corrupt an overlapping download' '
 	git init packfileclient-overlap &&
 	blob=$(test-tool genrandom pack-overlap 2m |
-- 
2.55.0.openai.131.g83a728de1eb6
Ted NymanJul 27, 2026, 00:28 UTC in reply to Ted Nyman on lore

[PATCH v6 6/6] fetch-pack: accept "pack" output for packfile URIs

When index-pack finds an existing keep file it reports pack rather than keep. Accept either result from http-fetch, and only register a keep lockfile when this fetch created it.

Read the pack/keep prefix and hash without consuming any following fsck output, validate the reported pack hash against the advertised hash, and exercise a packfile URI fetch with a pre-existing keep file.

Signed-off-by: Ted Nyman <tnyman@openai.com>
---
 fetch-pack.c           | 33 ++++++++++++++++++---------------
 t/t5702-protocol-v2.sh | 31 +++++++++++++++++++++++++++++++
 2 files changed, 49 insertions(+), 15 deletions(-)
Show changes to 2 files +49 −15

fetch-pack.c, t/t5702-protocol-v2.sh

diff --git a/fetch-pack.c b/fetch-pack.c
index 29c41132ee..e9f24fbd63 100644
--- a/fetch-pack.c
+++ b/fetch-pack.c
@@ -1887,9 +1887,10 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 	}
 
 	for (i = 0; i < packfile_uris.nr; i++) {
+		bool created_keep;
 		int j;
 		struct child_process cmd = CHILD_PROCESS_INIT;
-		char packname[GIT_MAX_HEXSZ + 1];
+		char packhash[GIT_MAX_HEXSZ + 1];
 		const char *uri = packfile_uris.items[i].string +
 			the_hash_algo->hexsz + 1;
 
@@ -1907,16 +1908,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (start_command(&cmd))
 			die("fetch-pack: unable to spawn http-fetch");
 
-		if (read_in_full(cmd.out, packname, 5) < 0 ||
-		    memcmp(packname, "keep\t", 5))
-			die("fetch-pack: expected keep then TAB at start of http-fetch output");
+		if (read_in_full(cmd.out, packhash, 5) != 5 ||
+		    (memcmp(packhash, "keep\t", 5) &&
+		     memcmp(packhash, "pack\t", 5)))
+			die("fetch-pack: expected pack or keep then TAB at start of http-fetch output");
+		created_keep = !memcmp(packhash, "keep\t", 5);
 
-		if (read_in_full(cmd.out, packname,
-				 the_hash_algo->hexsz + 1) < 0 ||
-		    packname[the_hash_algo->hexsz] != '\n')
-			die("fetch-pack: expected hash then LF at end of http-fetch output");
-
-		packname[the_hash_algo->hexsz] = '\0';
+		if (read_in_full(cmd.out, packhash,
+				 the_hash_algo->hexsz + 1) != the_hash_algo->hexsz + 1 ||
+		    packhash[the_hash_algo->hexsz] != '\n')
+			die("fetch-pack: expected hash then LF in http-fetch output");
+		packhash[the_hash_algo->hexsz] = '\0';
 
 		parse_gitmodules_oids(cmd.out, &fsck_options.gitmodules_found);
 
@@ -1925,16 +1927,17 @@ static struct ref *do_fetch_pack_v2(struct fetch_pack_args *args,
 		if (finish_command(&cmd))
 			die("fetch-pack: unable to finish http-fetch");
 
-		if (memcmp(packfile_uris.items[i].string, packname,
+		if (memcmp(packfile_uris.items[i].string, packhash,
 			   the_hash_algo->hexsz))
 			die("fetch-pack: pack downloaded from %s does not match expected hash %.*s",
 			    uri, (int) the_hash_algo->hexsz,
 			    packfile_uris.items[i].string);
 
-		string_list_append_nodup(pack_lockfiles,
-					 xstrfmt("%s/pack/pack-%s.keep",
-						 repo_get_object_directory(the_repository),
-						 packname));
+		if (created_keep)
+			string_list_append_nodup(pack_lockfiles,
+						 xstrfmt("%s/pack/pack-%s.keep",
+							 repo_get_object_directory(the_repository),
+							 packhash));
 	}
 	string_list_clear(&packfile_uris, 0);
 	strvec_clear(&index_pack_args);
diff --git a/t/t5702-protocol-v2.sh b/t/t5702-protocol-v2.sh
index 74a2b7730b..0f05286de8 100755
--- a/t/t5702-protocol-v2.sh
+++ b/t/t5702-protocol-v2.sh
@@ -1291,6 +1291,37 @@ test_expect_success 'packfile URIs with fetch instead of clone' '
 		fetch "$HTTPD_URL/smart/http_parent"
 '
 
+test_expect_success 'packfile URI preserves an existing keep file' '
+	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
+	rm -rf "$P" http_child keep.expect &&
+
+	git init "$P" &&
+	git -C "$P" config uploadpack.allowsidebandall true &&
+
+	echo my-blob >"$P/my-blob" &&
+	git -C "$P" add my-blob &&
+	git -C "$P" commit -m x &&
+	configure_exclusion "$P" my-blob >h &&
+
+	git init http_child &&
+	packhash=$(cat packh) &&
+	keep="http_child/.git/objects/pack/pack-$packhash.keep" &&
+	echo pre-existing >"$keep" &&
+	cp "$keep" keep.expect &&
+
+	GIT_TEST_SIDEBAND_ALL=1 \
+	git -C http_child -c protocol.version=2 \
+		-c fetch.uriprotocols=http,https \
+		fetch "$HTTPD_URL/smart/http_parent" &&
+
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.pack" &&
+	test_path_is_file \
+		"http_child/.git/objects/pack/pack-$packhash.idx" &&
+	test_cmp keep.expect "$keep" &&
+	git -C http_child cat-file -e "$(cat h)"
+'
+
 test_expect_success 'fetching with valid packfile URI but invalid hash fails' '
 	P="$HTTPD_DOCUMENT_ROOT_PATH/http_parent" &&
 	rm -rf "$P" http_child log &&
-- 
2.55.0.openai.131.g83a728de1eb6
Junio C HamanoJul 29, 2026, 21:41 UTC in reply to Ted Nyman on lore

Re: [PATCH v6 0/6] packfile URIs: support concurrent downloads

Ted Nyman <tnyman@openai.com> writes:
Show 17 quoted lines
> Changes since v5:
>
> * Split the existing double-close fix, HTTP 416 handling, generic
>   concurrent-download fix, and Windows sharing fix into separate
>   patches.
> * Replace the FIFO-based concurrent HTTP 416 test with a standalone
>   completed-partial test. Besides simplifying the test, this covers the
>   non-concurrent interrupted-download case directly.
> * Keep the final production code unchanged.
>
> Each patch passes t5550-http-fetch-dumb.sh. The final series also passes
> t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs
> with 12 parallel stress jobs.
>
> The v5 discussion is at:
>
> https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/
Is everybody happy with this new iteration?

The design of the re-download feature itself, as far as I understand, was favourably accepted from the earliest iteration, and now the CI breakages were corrected with the latest iteration of the tests, so we should be in pretty good shape, I presume.

Thanks.
Jeff KingAug 1, 2026, 13:53 UTC in reply to Ted Nyman on lore

Re: [PATCH v6 2/6] http: avoid closing index-pack input twice

On Sun, Jul 26, 2026 at 05:28:39PM -0700, Ted Nyman wrote:
Show 6 quoted lines
> finish_http_pack_request() passes its staging-file descriptor to
> index-pack through child_process.in. start_command() takes ownership
> of a supplied descriptor and closes it, even when starting the child
> fails.
> 
> Do not close the descriptor again after run_command() returns.
Thanks for splitting this out.
Show 12 quoted lines
> @@ -2704,13 +2704,8 @@ int finish_http_pack_request(struct http_pack_request *preq)
>  	else
>  		ip.no_stdout = 1;
>  
> -	if (run_command(&ip)) {
> +	if (run_command(&ip))
>  		ret = -1;
> -		goto cleanup;
> -	}
> -
> -cleanup:
> -	close(tmpfile_fd);

The patch _could_ just be a one-liner dropping this close(). Removing the cleanup label here is optional, but is a simplification that works because nobody else jumps to it (which must be true because we'd fail to compile otherwise).

I probably would have mentioned that in the commit message, but I think there's diminishing returns in trying to polish further.

-Peff
Jeff KingAug 1, 2026, 13:58 UTC in reply to Ted Nyman on lore

Re: [PATCH v6 3/6] http: accept HTTP 416 for complete partial packs

On Sun, Jul 26, 2026 at 05:28:40PM -0700, Ted Nyman wrote:
Show 11 quoted lines
> A resumed pack request may already have all bytes of the remote pack.
> A server can respond to the resulting Range request with HTTP 416
> instead of returning an empty response.
> 
> Accept that response in each pack-download caller and let index-pack
> validate the completed staging file. This can happen without concurrent
> downloads when a previous attempt completed the transfer but failed
> before indexing it.
> 
> Add a regression test that seeds a complete partial pack and checks that
> http-fetch indexes it after the server returns HTTP 416.

Again, thanks for splitting this out and demonstrating the non-concurrent case. It all looks good to me.

I do wonder what will happen when we get a 416 and we _don't_ have a complete pack. E.g., imagine the file size on the server changed (it shouldn't if they are using the hash of the pack contents as the name, but that's not strictly required).

Previously we'd barf on the curl error. Now we'll guess that we got the full file, even though we have a partial download. Presumably we'd then just barf at the index-pack level. I guess this is not really any different than other resumption problems. If the file changed on the server, we could easily download half of one version and half of another. Ultimately we don't trust any of it until index-pack processes the whole thing.

So this seems like a good direction to me.
-Peff
Jeff KingAug 1, 2026, 14:02 UTC in reply to Junio C Hamano on lore

Re: [PATCH v6 0/6] packfile URIs: support concurrent downloads

On Wed, Jul 29, 2026 at 02:41:51PM -0700, Junio C Hamano wrote:
Show 26 quoted lines
> Ted Nyman <tnyman@openai.com> writes:
> 
> > Changes since v5:
> >
> > * Split the existing double-close fix, HTTP 416 handling, generic
> >   concurrent-download fix, and Windows sharing fix into separate
> >   patches.
> > * Replace the FIFO-based concurrent HTTP 416 test with a standalone
> >   completed-partial test. Besides simplifying the test, this covers the
> >   non-concurrent interrupted-download case directly.
> > * Keep the final production code unchanged.
> >
> > Each patch passes t5550-http-fetch-dumb.sh. The final series also passes
> > t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs
> > with 12 parallel stress jobs.
> >
> > The v5 discussion is at:
> >
> > https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/
> 
> Is everybody happy with this new iteration?
> 
> The design of the re-download feature itself, as far as I
> understand, was favourably accepted from the earliest iteration, and
> now the CI breakages were corrected with the latest iteration of the
> tests, so we should be in pretty good shape, I presume.

Yeah, sorry, I hadn't had time to look carefully. I just did so, and it all looks good to me. v6 splits the patches in a way that (at least to my mind) make the trickiest parts of the logic easier to follow.

-Peff
Junio C HamanoAug 8, 2026, 16:23 UTC in reply to Jeff King on lore

Re: [PATCH v6 0/6] packfile URIs: support concurrent downloads

Jeff King <peff@peff.net> writes:
Show 32 quoted lines
> On Wed, Jul 29, 2026 at 02:41:51PM -0700, Junio C Hamano wrote:
>
>> Ted Nyman <tnyman@openai.com> writes:
>> 
>> > Changes since v5:
>> >
>> > * Split the existing double-close fix, HTTP 416 handling, generic
>> >   concurrent-download fix, and Windows sharing fix into separate
>> >   patches.
>> > * Replace the FIFO-based concurrent HTTP 416 test with a standalone
>> >   completed-partial test. Besides simplifying the test, this covers the
>> >   non-concurrent interrupted-download case directly.
>> > * Keep the final production code unchanged.
>> >
>> > Each patch passes t5550-http-fetch-dumb.sh. The final series also passes
>> > t5702-protocol-v2.sh, and the overlapping-download test passes 240 runs
>> > with 12 parallel stress jobs.
>> >
>> > The v5 discussion is at:
>> >
>> > https://lore.kernel.org/git/cover.1785047139.git.tnyman@openai.com/
>> 
>> Is everybody happy with this new iteration?
>> 
>> The design of the re-download feature itself, as far as I
>> understand, was favourably accepted from the earliest iteration, and
>> now the CI breakages were corrected with the latest iteration of the
>> tests, so we should be in pretty good shape, I presume.
>
> Yeah, sorry, I hadn't had time to look carefully. I just did so, and it
> all looks good to me. v6 splits the patches in a way that (at least to
> my mind) make the trickiest parts of the logic easier to follow.
Thanks.

Back to recent threads