git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v4 0/3] packfile URIs: support concurrent downloads

From
Ted Nyman <tnyman@openai.com>
Date
Jul 24, 2026, 08:14 UTC
Message-ID
<cover.1784874850.git.tnyman@openai.com>
In-Reply-To
<cover.1784676106.git.tnyman@openai.com>

Packfile URI and dumb HTTP downloads stage packs at objects/pack/pack-<hash>.pack.temp so an interrupted transfer can resume. Opening that file in append mode forces every write to its current end. Two Git processes fetching the same pack into one object database can therefore append duplicate data and corrupt the pack.

The first patch separates the unrelated --index-pack-arg documentation and error-message correction requested during review.

The second patch keeps the predictable staging name but removes append mode. Each downloader seeks once to the current end, requests the corresponding Range, and writes using its own descriptor offset. Since the staging key must identify immutable pack contents, overlapping responses write identical bytes at identical offsets. There is no need for pwrite(2) or cross-process coordination, and resumption continues to work for both packfile URI and ordinary dumb HTTP downloads.

A downloader can also find that the partial pack has completed and request a range starting at EOF. Servers may respond with HTTP 416 in that case. Treat the response as a completed download and let index-pack validate the pack.

On MinGW, the non-append O_RDWR open grants FILE_SHARE_DELETE only for an existing file. Create a missing staging file exclusively, close it, and reopen it without O_CREAT so every retained descriptor permits another downloader to unlink the path. Keep the open descriptor for index-pack; it installs its own pack, so the shared staging file is only unlinked, never renamed.

The third patch handles the related .keep race. When another process has already created the keep file, index-pack reports "pack<TAB><hash>" instead of "keep<TAB><hash>". Accept both successful forms and remove only keep files created by the current process. Read only the prefix and hash so any following fsck output remains available to fetch-pack.

The tests cover resumption, a completed partial returning 416, overlapping 200 and 206 responses, unlinking the staging path while index-pack holds its descriptor, and a pre-existing .keep file. The unlink test does not require FIFOs, so it can exercise MinGW's sharing behavior even though the concurrent-download tests are skipped there.

Changes since v3:
  * Match HTTP 416 in trace output from both older and current libcurl.
  * Add a timeout to the overlapping-download test server, notify FIFO
    waiters on server failures, and track the actual server process for
    cleanup.
  * Wait for the second downloader first so an early failure cannot
    leave the test server waiting for a request that will never arrive.
  * No production code changes.

These changes avoid false failures with older libcurl and prevent a failed downloader from leaving the test server running indefinitely.

The v3 discussion is at:
  https://lore.kernel.org/git/cover.1784676106.git.tnyman@openai.com/
Ted Nyman (3):
  http-fetch: correct --index-pack-arg documentation
  http: avoid concurrent appends to partial packs
  fetch-pack: accept "pack" output for packfile URIs
 Documentation/git-http-fetch.adoc |  13 +-
 fetch-pack.c                      |  33 ++--
 http-fetch.c                      |   7 +-
 http-push.c                       |   3 +-
 http-walker.c                     |   3 +-
 http.c                            |  56 ++++---
 t/t5550-http-fetch-dumb.sh        | 250 ++++++++++++++++++++++++++++++
 t/t5702-protocol-v2.sh            |  31 ++++
 8 files changed, 350 insertions(+), 46 deletions(-)
Range-diff against v3:
1:  a6a40b8046 = 1:  a6a40b8046 http-fetch: correct --index-pack-arg documentation
2:  6c91054afc ! 2:  144c98cdfa http: avoid concurrent appends to partial packs
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	test_cmp expect second.out &&
     +	test_grep "Range: bytes=64-" first.trace &&
     +	test_grep "Range: bytes=[0-9]*-" second.trace &&
    -+	test_grep "HTTP/[0-9.]* 416" second.trace &&
    ++	test_grep "416 Requested Range Not Satisfiable" second.trace &&
     +	test_path_is_missing "$tmpfile" &&
     +	git -C packfileclient-concurrent cat-file -e "$HASH"
     +'
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	use IO::Socket::INET;
     +
     +	my ($packfile, $server_ready, $first_ready) = @ARGV;
    ++	my $completed = 0;
    ++	END {
    ++		if (!$completed) {
    ++			signal_ready($server_ready, "failed");
    ++			signal_ready($first_ready, "failed");
    ++		}
    ++	}
    ++
    ++	$SIG{ALRM} = sub { die "timed out serving concurrent pack requests\n" };
    ++	alarm 60;
    ++
     +	open(my $in, "<:raw", $packfile) or die "open $packfile: $!";
     +	my $pack = do { local $/; <$in> };
     +	close($in) or die "close $packfile: $!";
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +	write_all($second, substr($pack, $second_pos));
     +	close($first) or die "close first response: $!";
     +	close($second) or die "close second response: $!";
    ++	$completed = 1;
    ++	alarm 0;
     +	EOF
     +	{
    -+		(
    -+			if ! "$TRASH_DIRECTORY/slow-pack-server" "$pack" \
    -+				"$TRASH_DIRECTORY/server-ready" \
    -+				"$TRASH_DIRECTORY/first-ready"
    -+			then
    -+				echo failed >"$TRASH_DIRECTORY/server-ready" &&
    -+				echo failed >"$TRASH_DIRECTORY/first-ready" &&
    -+				exit 1
    -+			fi
    -+		) >server.log 2>&1 &
    ++		"$TRASH_DIRECTORY/slow-pack-server" "$pack" \
    ++			"$TRASH_DIRECTORY/server-ready" \
    ++			"$TRASH_DIRECTORY/first-ready" >server.log 2>&1 &
     +		server_pid=$!
     +	} &&
     +	test_when_finished "
    @@ t/t5550-http-fetch-dumb.sh: test_expect_success 'http-fetch --packfile' '
     +		kill $second_pid 2>/dev/null || :
     +		wait $second_pid 2>/dev/null || :
     +	" &&
    -+	wait "$server_pid" &&
    -+	wait "$first_pid" &&
     +	wait "$second_pid" &&
    ++	wait "$first_pid" &&
    ++	wait "$server_pid" &&
     +	test_grep "HTTP/[0-9.]* 200" overlap-first.trace &&
     +	test_grep "Range: bytes=[1-9][0-9]*-" overlap-second.trace &&
     +	test_grep "HTTP/[0-9.]* 206" overlap-second.trace &&
3:  1ee5d7e027 = 3:  d9063deb60 fetch-pack: accept "pack" output for packfile URIs
base-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200
-- 
2.55.0.openai.131.g83a728de1eb6
Previous: Junio C HamanoNext: Ted Nyman
Message 32 of 57 in “packfile URIs: support concurrent downloads”
  1. 0/2 packfile URIs: support concurrent downloadsTed Nyman, Jul 13, 2026
  2. 1/2 http: use unique tempfiles for packfile URI downloadsTed Nyman, Jul 13, 2026
  3. Junio C HamanoJul 14, 2026
  4. Ted NymanJul 14, 2026
  5. Taylor BlauJul 14, 2026
  6. Jeff KingJul 14, 2026
  7. Junio C HamanoJul 14, 2026
  8. Ted NymanJul 14, 2026
  9. Taylor BlauJul 14, 2026
  10. Jeff KingJul 14, 2026
  11. Jeff KingJul 14, 2026
  12. 2/2 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 13, 2026
  13. Jeff KingJul 14, 2026
  14. Jeff KingJul 14, 2026
  15. Ted NymanJul 14, 2026
  16. Jeff KingJul 14, 2026
  17. Taylor BlauJul 14, 2026
  18. 0/2 packfile URIs: support concurrent downloadsTed Nyman, Jul 20, 2026
  19. 1/2 http: avoid concurrent appends to partial packsTed Nyman, Jul 20, 2026
  20. Junio C HamanoJul 21, 2026
  21. 2/2 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 20, 2026
  22. 0/3 packfile URIs: support concurrent downloadsTed Nyman, Jul 21, 2026
  23. 1/3 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 21, 2026
  24. 2/3 http: avoid concurrent appends to partial packsTed Nyman, Jul 21, 2026
  25. 3/3 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 21, 2026
  26. Junio C HamanoJul 24, 2026
  27. Jeff KingJul 25, 2026
  28. Jeff KingJul 25, 2026
  29. Jeff KingJul 25, 2026
  30. Jeff KingJul 25, 2026
  31. Junio C HamanoJul 25, 2026
  32. 0/3 packfile URIs: support concurrent downloadsTed Nyman, Jul 24, 2026
  33. 1/3 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 24, 2026
  34. Taylor BlauJul 24, 2026
  35. 2/3 http: avoid concurrent appends to partial packsTed Nyman, Jul 24, 2026
  36. 3/3 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 24, 2026
  37. Taylor BlauJul 24, 2026
  38. 0/3 packfile URIs: support concurrent downloadsTed Nyman, Jul 26, 2026
  39. 1/3 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 26, 2026
  40. 2/3 http: avoid concurrent appends to partial packsTed Nyman, Jul 26, 2026
  41. Jeff KingJul 26, 2026
  42. Ted NymanJul 26, 2026
  43. Jeff KingJul 26, 2026
  44. 3/3 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 26, 2026
  45. Jeff KingJul 26, 2026
  46. 0/6 packfile URIs: support concurrent downloadsTed Nyman, Jul 27, 2026
  47. 1/6 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 27, 2026
  48. 2/6 http: avoid closing index-pack input twiceTed Nyman, Jul 27, 2026
  49. Jeff KingAug 1, 2026
  50. 3/6 http: accept HTTP 416 for complete partial packsTed Nyman, Jul 27, 2026
  51. Jeff KingAug 1, 2026
  52. 4/6 http: avoid concurrent appends to partial packsTed Nyman, Jul 27, 2026
  53. 5/6 http: permit unlinking partial packs on WindowsTed Nyman, Jul 27, 2026
  54. 6/6 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 27, 2026
  55. Junio C HamanoJul 29, 2026
  56. Jeff KingAug 1, 2026
  57. Junio C HamanoAug 8, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.