git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v2 1/2] http: avoid concurrent appends to partial packs

From
Junio C Hamano <gitster@pobox.com>
Date
Jul 21, 2026, 19:56 UTC
Message-ID
<xmqqo6g0unfb.fsf@gitster.g>
In-Reply-To
<160a9b9fd0982dadfbf6f8fbb378d1a3e9173698.1784582665.git.tnyman@openai.com>
Ted Nyman <tnyman@openai.com> writes:
Show 16 quoted lines
> Pack requests stage downloads in a predictable partial-pack file so an
> interrupted transfer can be resumed. Both packfile URI and ordinary dumb
> HTTP requests use this staging path. Opening it in append mode lets
> concurrent fetches interleave their writes, corrupting the pack or
> causing a later fetch to request a range at EOF.
>
> Open the partial pack read-write, seek to its current end, and retain a
> per-descriptor offset for incoming data. Reopen newly created partial
> packs without O_CREAT so Windows permits concurrent unlink, and keep the
> descriptor for index-pack when another downloader removes the staging
> path. Accept HTTP 416 when a partial pack is already complete.
>
> Exercise resumed transfers, EOF ranges, and overlapping 200 and 206
> responses. Clarify the staging-key documentation and correct the stale
> --index-pack-args spelling in the documentation and error messages; the
> repeatable --index-pack-arg option is already accepted.

Hmph. So the idea is to allow multiple processes to open the same file and, because they all know where their respective chunks of data fit in the final file, have them use pwrite(2) to deposit those pieces at the exact target locations, and this prevents them from stepping on each other's toes?

I cannot exactly explain why but it somehow makes me feel dirty.

It is also surprising that the workaround on MinGW works when one of these multiple processes finishes writing and attempts to finalize the temporary file while others still have open file descriptors to the same file.

Show 7 quoted lines
> -	The hash is used to determine the name of the temporary file and is
> -	arbitrary. The output of index-pack is printed to stdout. Requires
> -	--index-pack-args.
> +	The hash is used to determine the name of the temporary file. It need
> +	not be the pack hash, but it must uniquely identify the pack contents
> +	for resumption. The output of index-pack is printed to stdout. Requires
> +	one or more --index-pack-arg options.
OK.
Show 6 quoted lines
> ---index-pack-args=<args>::
> -	For internal use only. The command to run on the contents of the
> -	downloaded pack. Arguments are URL-encoded separated by spaces.
> +--index-pack-arg=<arg>::
> +	For internal use only. An argument to the command run on the contents
> +	of the downloaded pack. This option can be specified multiple times.

Was the 'internal use only' thing renamed in order to prevent the new code from accidentally working with an older caller?

    ... goes and notices that the code uses singular form throughout ...

Ah, no, this is an unrelated typo fix that remains valid even if the rest of this patch is dropped. Good catch.

It would be easier to review the actual changes if this cleanup were isolated in a preliminary patch. Are there other cleanup changes in this series that fall into the same category?

Show 11 quoted lines
> diff --git a/http-fetch.c b/http-fetch.c
> index f9b6ecb061..05f68f306a 100644
> --- a/http-fetch.c
> +++ b/http-fetch.c
> @@ -70,7 +70,8 @@ static void fetch_single_packfile(struct object_id *packfile_hash,
>  
>  	if (start_active_slot(preq->slot)) {
>  		run_active_slot(preq->slot);
> -		if (results.curl_result != CURLE_OK) {
> +		if (results.curl_result != CURLE_OK &&
> +		    results.http_code != 416) {

We do not seem to use symbolic constants for these '4xx' codes (or '2xx', for that matter), so I will let that pass. Eventually, we may want to give symbolic constants to them to improve readability, but doing so is certainly outside the scope of this topic.

Show 6 quoted lines
> @@ -155,7 +156,7 @@ int cmd_main(int argc, const char **argv)
>  
>  	if (packfile) {
>  		if (!index_pack_args.nr)
> -			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-args");
> +			die(_("the option '%s' requires '%s'"), "--packfile", "--index-pack-arg");
This and ...
Show 9 quoted lines
> @@ -164,7 +165,7 @@ int cmd_main(int argc, const char **argv)
>  	}
>  
>  	if (index_pack_args.nr)
> -		die(_("the option '%s' requires '%s'"), "--index-pack-args", "--packfile");
> +		die(_("the option '%s' requires '%s'"), "--index-pack-arg", "--packfile");
>  
>  	if (commits_on_stdin) {
>  		commits = walker_targets_stdin(&commit_id, &write_ref);

... this is the same "index-pack-arg" fix and can be moved to a separate preliminary clean-up patch.

Thanks.
Previous: Ted NymanNext: Ted Nyman
Message 20 of 57 in “packfile URIs: support concurrent downloads”
  1. 0/2 packfile URIs: support concurrent downloadsTed Nyman, Jul 13, 2026
  2. 1/2 http: use unique tempfiles for packfile URI downloadsTed Nyman, Jul 13, 2026
  3. Junio C HamanoJul 14, 2026
  4. Ted NymanJul 14, 2026
  5. Taylor BlauJul 14, 2026
  6. Jeff KingJul 14, 2026
  7. Junio C HamanoJul 14, 2026
  8. Ted NymanJul 14, 2026
  9. Taylor BlauJul 14, 2026
  10. Jeff KingJul 14, 2026
  11. Jeff KingJul 14, 2026
  12. 2/2 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 13, 2026
  13. Jeff KingJul 14, 2026
  14. Jeff KingJul 14, 2026
  15. Ted NymanJul 14, 2026
  16. Jeff KingJul 14, 2026
  17. Taylor BlauJul 14, 2026
  18. 0/2 packfile URIs: support concurrent downloadsTed Nyman, Jul 20, 2026
  19. 1/2 http: avoid concurrent appends to partial packsTed Nyman, Jul 20, 2026
  20. Junio C HamanoJul 21, 2026
  21. 2/2 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 20, 2026
  22. 0/3 packfile URIs: support concurrent downloadsTed Nyman, Jul 21, 2026
  23. 1/3 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 21, 2026
  24. 2/3 http: avoid concurrent appends to partial packsTed Nyman, Jul 21, 2026
  25. 3/3 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 21, 2026
  26. Junio C HamanoJul 24, 2026
  27. Jeff KingJul 25, 2026
  28. Jeff KingJul 25, 2026
  29. Jeff KingJul 25, 2026
  30. Jeff KingJul 25, 2026
  31. Junio C HamanoJul 25, 2026
  32. 0/3 packfile URIs: support concurrent downloadsTed Nyman, Jul 24, 2026
  33. 1/3 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 24, 2026
  34. Taylor BlauJul 24, 2026
  35. 2/3 http: avoid concurrent appends to partial packsTed Nyman, Jul 24, 2026
  36. 3/3 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 24, 2026
  37. Taylor BlauJul 24, 2026
  38. 0/3 packfile URIs: support concurrent downloadsTed Nyman, Jul 26, 2026
  39. 1/3 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 26, 2026
  40. 2/3 http: avoid concurrent appends to partial packsTed Nyman, Jul 26, 2026
  41. Jeff KingJul 26, 2026
  42. Ted NymanJul 26, 2026
  43. Jeff KingJul 26, 2026
  44. 3/3 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 26, 2026
  45. Jeff KingJul 26, 2026
  46. 0/6 packfile URIs: support concurrent downloadsTed Nyman, Jul 27, 2026
  47. 1/6 http-fetch: correct --index-pack-arg documentationTed Nyman, Jul 27, 2026
  48. 2/6 http: avoid closing index-pack input twiceTed Nyman, Jul 27, 2026
  49. Jeff KingAug 1, 2026
  50. 3/6 http: accept HTTP 416 for complete partial packsTed Nyman, Jul 27, 2026
  51. Jeff KingAug 1, 2026
  52. 4/6 http: avoid concurrent appends to partial packsTed Nyman, Jul 27, 2026
  53. 5/6 http: permit unlinking partial packs on WindowsTed Nyman, Jul 27, 2026
  54. 6/6 fetch-pack: accept "pack" output for packfile URIsTed Nyman, Jul 27, 2026
  55. Junio C HamanoJul 29, 2026
  56. Jeff KingAug 1, 2026
  57. Junio C HamanoAug 8, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.