Volume XXII, number 279Tuesday, October 6, 2026Latest message 35 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

patch, 2 partsci: reduce pressure from large test fixtures

25 messages between Sep 23, 2026 and Sep 30, 2026, from Tamir Duberstein, Patrick Steinhardt, Jeff King, Junio C Hamano.

Plain Markdown or JSON for tools and agents. Diffs are folded; open one to read it.

Tamir DubersteinSep 23, 2026, 17:13 UTC on lore

Linux pull-request jobs with the long tests enabled hit two resource problems: diff was killed after git log produced its huge output, and large clone and repack tests hit ENOSPC.

The first patch compares the huge output with cmp and removes the large files after success. The second sizes Make and prove parallelism to the Linux runner's available CPUs, following the existing GitLab policy. Both changes keep the large test cases enabled.

Prepared with Codex, including review by a separate Codex agent.
Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
Tamir Duberstein (2):
      t4205: compare huge output without diff
      ci: match Linux jobs to available CPUs
 ci/lib.sh                     | 4 ++++
 t/t4205-log-pretty-formats.sh | 3 ++-
 2 files changed, 6 insertions(+), 1 deletion(-)

--- base-commit: 3bc0341126508f78f5869cbfc0005e987efdf0c7 change-id: 20260923-ci-large-test-resources-349cdc95f7f8

Tamir DubersteinSep 23, 2026, 17:13 UTC in reply to Tamir Duberstein on lore

[PATCH 1/2] t4205: compare huge output without diff

The huge-commit test compares two files with a line larger than 2 GiB. In Linux GitHub Actions jobs, git log produces its huge output but its subsequent diff process is killed with SIGKILL.

Use test_cmp_bin to compare the output byte for byte without constructing a line-oriented diff. Remove the two large files after a successful comparison, releasing more than 4 GiB before subsequent tests.

Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
 t/t4205-log-pretty-formats.sh | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)
Show changes to t/t4205-log-pretty-formats.sh +2 −1
diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
index 4be5c51489..6279a7e9bc 100755
--- a/t/t4205-log-pretty-formats.sh
+++ b/t/t4205-log-pretty-formats.sh
@@ -1189,7 +1189,8 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
 test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
 	git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
 	echo 0 >>expect &&
-	test_cmp expect actual
+	test_cmp_bin expect actual &&
+	rm expect actual
 '
 
 test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message does not cause allocation failure' '
-- 
2.56.0.rc0.807.ga0c0929ce1.frankengit
Tamir DubersteinSep 23, 2026, 17:13 UTC in reply to Tamir Duberstein on lore

[PATCH 2/2] ci: match Linux jobs to available CPUs

GitHub Actions runs ten Make test suites concurrently even on private Linux runners with two CPUs. Pull request runs enable the long tests. These runs hit ENOSPC while multiple multi-gigabyte clone and repack fixtures were active.

Use nproc to choose Make and prove parallelism, as the GitLab CI path already does. This reduces overlapping fixtures on small Linux runners while keeping the long tests enabled.

Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
 ci/lib.sh | 4 ++++
 1 file changed, 4 insertions(+)
Show changes to ci/lib.sh +4 −0
diff --git a/ci/lib.sh b/ci/lib.sh
index c6ccbf8c17..0855026dad 100755
--- a/ci/lib.sh
+++ b/ci/lib.sh
@@ -228,6 +228,10 @@ then
 
 	GIT_TEST_OPTS="--github-workflow-markup"
 	JOBS=10
+	if test linux = "$CI_OS_NAME"
+	then
+		JOBS=$(nproc)
+	fi
 
 	distro=$(echo "$CI_JOB_IMAGE" | tr : -)
 elif test true = "$GITLAB_CI"
-- 
2.56.0.rc0.807.ga0c0929ce1.frankengit
Tamir DubersteinSep 23, 2026, 18:36 UTC in reply to Tamir Duberstein on lore

Re: [PATCH 0/2] ci: reduce pressure from large test fixtures

On Wed, Sep 23, 2026 at 1:13 PM Tamir Duberstein <tamird@gmail.com> wrote:
>
> Prepared with Codex, including review by a separate Codex agent.
Apologies for this inclusion, won't happen again.
Patrick SteinhardtSep 24, 2026, 06:13 UTC in reply to Tamir Duberstein on lore

Re: [PATCH 1/2] t4205: compare huge output without diff

On Wed, Sep 23, 2026 at 01:13:28PM -0400, Tamir Duberstein wrote:
> The huge-commit test compares two files with a line larger than 2 GiB.
> In Linux GitHub Actions jobs, git log produces its huge output but
> its subsequent diff process is killed with SIGKILL.
I've never seen that failure before. Do you maybe have a link to it?
> Use test_cmp_bin to compare the output byte for byte without constructing
> a line-oriented diff. Remove the two large files after a successful
> comparison, releasing more than 4 GiB before subsequent tests.

It would be great to back up the claim that test_cmp_bin is better than test_cmp, e.g. by comparing peak RSS and its runtime.

Show 17 quoted lines
> Signed-off-by: Tamir Duberstein <tamird@gmail.com>
> ---
>  t/t4205-log-pretty-formats.sh | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
> 
> diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
> index 4be5c51489..6279a7e9bc 100755
> --- a/t/t4205-log-pretty-formats.sh
> +++ b/t/t4205-log-pretty-formats.sh
> @@ -1189,7 +1189,8 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
>  test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
>  	git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
>  	echo 0 >>expect &&
> -	test_cmp expect actual
> +	test_cmp_bin expect actual &&
> +	rm expect actual
>  '

Hm. Sure, releasing these files isn't a bad idea by itself. But we rewrite "expect" in the next test anyway, and "actual" will be rewritten two tests further down. So does it really buy us that much...?

Patrick
Patrick SteinhardtSep 24, 2026, 06:13 UTC in reply to Tamir Duberstein on lore

Re: [PATCH 2/2] ci: match Linux jobs to available CPUs

On Wed, Sep 23, 2026 at 01:13:29PM -0400, Tamir Duberstein wrote:
> GitHub Actions runs ten Make test suites concurrently even on private
> Linux runners with two CPUs. Pull request runs enable the long tests.
> These runs hit ENOSPC while multiple multi-gigabyte clone and repack
> fixtures were active.
Again, a link would be appreciated that demonstrates this.
> Use nproc to choose Make and prove parallelism, as the GitLab CI path
> already does. This reduces overlapping fixtures on small Linux runners
> while keeping the long tests enabled.

It may avoid overlapping fixtures. But what does CI runtime look like before and after this change? Does it improve? Does it regress? Would it maybe make sense to oversubscribe at least a bit?

Show 12 quoted lines
> diff --git a/ci/lib.sh b/ci/lib.sh
> index c6ccbf8c17..0855026dad 100755
> --- a/ci/lib.sh
> +++ b/ci/lib.sh
> @@ -228,6 +228,10 @@ then
>  
>  	GIT_TEST_OPTS="--github-workflow-markup"
>  	JOBS=10
> +	if test linux = "$CI_OS_NAME"
> +	then
> +		JOBS=$(nproc)
> +	fi

Makes me wonder whether we should have the same logic on both GitLab and GitHub going forward. There probably isn't a good reason why these two should differ from one another.

Patrick
Tamir DubersteinSep 24, 2026, 20:41 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH 1/2] t4205: compare huge output without diff

On Thu, Sep 24, 2026 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:
Show 7 quoted lines
>
> On Wed, Sep 23, 2026 at 01:13:28PM -0400, Tamir Duberstein wrote:
> > The huge-commit test compares two files with a line larger than 2 GiB.
> > In Linux GitHub Actions jobs, git log produces its huge output but
> > its subsequent diff process is killed with SIGKILL.
>
> I've never seen that failure before. Do you maybe have a link to it?
The failures happened on a private repo that I've since lost access to
- but I believe it was precipitated by GitHub runners having half the
memory in private repos as in public ones [1].
Show 6 quoted lines
> > Use test_cmp_bin to compare the output byte for byte without constructing
> > a line-oriented diff. Remove the two large files after a successful
> > comparison, releasing more than 4 GiB before subsequent tests.
>
> It would be great to back up the claim that test_cmp_bin is better than
> test_cmp, e.g. by comparing peak RSS and its runtime.

As for the comparison: on Linux arm64 with GNU diffutils 3.8 using two identical files containing 2,147,483,649 "1" bytes followed by "0\n" (matching this test's expected output) gave:

Command Mean +/- stddev Maximum RSS (KiB) diff -u expect actual 5.276 +/- 0.572 s 4199924 cmp expect actual 0.506 +/- 0.099 s 1264

Show 22 quoted lines
>
> > Signed-off-by: Tamir Duberstein <tamird@gmail.com>
> > ---
> >  t/t4205-log-pretty-formats.sh | 3 ++-
> >  1 file changed, 2 insertions(+), 1 deletion(-)
> >
> > diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
> > index 4be5c51489..6279a7e9bc 100755
> > --- a/t/t4205-log-pretty-formats.sh
> > +++ b/t/t4205-log-pretty-formats.sh
> > @@ -1189,7 +1189,8 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
> >  test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
> >       git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
> >       echo 0 >>expect &&
> > -     test_cmp expect actual
> > +     test_cmp_bin expect actual &&
> > +     rm expect actual
> >  '
>
> Hm. Sure, releasing these files isn't a bad idea by itself. But we
> rewrite "expect" in the next test anyway, and "actual" will be rewritten
> two tests further down. So does it really buy us that much...?
You're right, this probably does not buy much.

Would you like me to include the performance comparison in v2? As for the deletion: would you prefer I drop it?

Link: https://docs.github.com/en/actions/reference/runners/github-hosted-runners#standard-github-hosted-runners-for--private-repositories
[1]
Jeff KingSep 25, 2026, 00:13 UTC in reply to Tamir Duberstein on lore

Re: [PATCH 1/2] t4205: compare huge output without diff

On Thu, Sep 24, 2026 at 04:41:05PM -0400, Tamir Duberstein wrote:
Show 10 quoted lines
> > It would be great to back up the claim that test_cmp_bin is better than
> > test_cmp, e.g. by comparing peak RSS and its runtime.
> 
> As for the comparison: on Linux arm64 with GNU
> diffutils 3.8 using two identical files containing 2,147,483,649 "1" bytes
> followed by "0\n" (matching this test's expected output) gave:
> 
> Command              Mean +/- stddev       Maximum RSS (KiB)
> diff -u expect actual  5.276 +/- 0.572 s              4199924
> cmp expect actual      0.506 +/- 0.099 s                 1264

Yeah, that's a big difference. This is probably an outlier because of the giant files, but I've wondered if test_cmp() ought to be doing something like:

  # first check quickly if there is any difference
  cmp "$@" && return 0
  # if not, then we can spend time to produce a useful output
  diff "$@"

But I never pursued it because I figured most of the things we compare are not big enough to matter (and really the comparison itself is dwarfed by the process startup time).

So I am happy just marking this particular case as cmp_bin, too.
-Peff
Junio C HamanoSep 25, 2026, 06:41 UTC in reply to Jeff King on lore

Re: [PATCH 1/2] t4205: compare huge output without diff

Jeff King <peff@peff.net> writes:
Show 8 quoted lines
>   # first check quickly if there is any difference
>   cmp "$@" && return 0
>   # if not, then we can spend time to produce a useful output
>   diff "$@"
>
> But I never pursued it because I figured most of the things we compare
> are not big enough to matter (and really the comparison itself is
> dwarfed by the process startup time).
Yup, that matches my intuition.
> So I am happy just marking this particular case as cmp_bin, too.
Me too.
Thanks.
Tamir DubersteinSep 25, 2026, 15:03 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH 2/2] ci: match Linux jobs to available CPUs

On Thu, Sep 24, 2026 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:
Show 8 quoted lines
>
> On Wed, Sep 23, 2026 at 01:13:29PM -0400, Tamir Duberstein wrote:
> > GitHub Actions runs ten Make test suites concurrently even on private
> > Linux runners with two CPUs. Pull request runs enable the long tests.
> > These runs hit ENOSPC while multiple multi-gigabyte clone and repack
> > fixtures were active.
>
> Again, a link would be appreciated that demonstrates this.

Yeah, sorry about that. Again, I suspect GitHub runners on private repos (2 CPUs) make this worse than public repos (4 CPUs).

Show 7 quoted lines
> > Use nproc to choose Make and prove parallelism, as the GitLab CI path
> > already does. This reduces overlapping fixtures on small Linux runners
> > while keeping the long tests enabled.
>
> It may avoid overlapping fixtures. But what does CI runtime look like
> before and after this change? Does it improve? Does it regress? Would it
> maybe make sense to oversubscribe at least a bit?

I ran a bunch of tests of this in GitHub and the disappointing answer is that it's not clear what the right choice is:

                    Fixed 10   CPU count   2x CPU count
Linux Make             278.9       273.9          259.9
macOS Make              94.5       119.1           99.8
Windows Make           102.2       103.8          100.8
Show 17 quoted lines
>
> > diff --git a/ci/lib.sh b/ci/lib.sh
> > index c6ccbf8c17..0855026dad 100755
> > --- a/ci/lib.sh
> > +++ b/ci/lib.sh
> > @@ -228,6 +228,10 @@ then
> >
> >       GIT_TEST_OPTS="--github-workflow-markup"
> >       JOBS=10
> > +     if test linux = "$CI_OS_NAME"
> > +     then
> > +             JOBS=$(nproc)
> > +     fi
>
> Makes me wonder whether we should have the same logic on both GitLab and
> GitHub going forward. There probably isn't a good reason why these two
> should differ from one another.

Agreed. I have rewritten this patch to use the same logic across CI providers (1x CPU count) in v2. I'll leave tuning (e.g. moving to 2x) to a future change.

>
> Patrick
Tamir DubersteinSep 25, 2026, 16:35 UTC in reply to Tamir Duberstein on lore

[PATCH v2 0/2] ci: use cmp and align job-count selection

The first patch uses cmp for the huge commit-message output comparison. On this input, GNU diffutils 3.8 on Linux arm64 takes 5.276 seconds with 4,199,924 KiB peak RSS for diff, versus 0.506 seconds and 1,264 KiB for cmp. Patch 1 includes the fixture and measurement details.

The second patch aligns GitHub Actions' Make and prove job counts with GitLab CI's CPU-count policy, replacing GitHub's fixed ten jobs.

Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
Changes in v2:
- Replace the unavailable CI failure reference with comparison runtime
  and peak RSS measurements.
- Drop the file removal; following tests overwrite expect and actual.
- Share job-count selection between GitHub Actions and GitLab CI.
- Use native CPU-count queries on macOS and Windows.
- Link to v1: https://patch.msgid.link/20260923-ci-large-test-resources-v1-0-c28416d59475@gmail.com
---
Tamir Duberstein (2):
      t4205: compare huge output without diff
      ci: align job counts across CI providers
 ci/lib.sh                     | 16 ++++++++++++----
 t/t4205-log-pretty-formats.sh |  2 +-
 2 files changed, 13 insertions(+), 5 deletions(-)
Range-diff versus v1:
1:  5babc36eb5 ! 1:  b48c86e164 t4205: compare huge output without diff
    @@ Metadata
      ## Commit message ##
         t4205: compare huge output without diff
     
    -    The huge-commit test compares two files with a line larger than 2 GiB.
    -    In Linux GitHub Actions jobs, git log produces its huge output but
    -    its subsequent diff process is killed with SIGKILL.
    +    The huge-commit test compares output containing a line larger than 2 GiB.
    +    For two identical files containing 2,147,483,649 "1" bytes followed by
    +    "0\n", GNU diffutils 3.8 on Linux arm64 gives these measurements:
     
    -    Use test_cmp_bin to compare the output byte for byte without constructing
    -    a line-oriented diff. Remove the two large files after a successful
    -    comparison, releasing more than 4 GiB before subsequent tests.
    +      Command               Mean +/- stddev       Maximum RSS (KiB)
    +      diff -u expect actual  5.276 +/- 0.572 s              4199924
    +      cmp expect actual      0.506 +/- 0.099 s                 1264
     
    +    The test needs only an equality check. Use test_cmp_bin, which runs cmp,
    +    to compare the output byte for byte with less time and memory.
    +
    +    Assisted-by: LLM
         Signed-off-by: Tamir Duberstein <tamird@gmail.com>
     
      ## t/t4205-log-pretty-formats.sh ##
    @@ t/t4205-log-pretty-formats.sh: test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'se
      	git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
      	echo 0 >>expect &&
     -	test_cmp expect actual
    -+	test_cmp_bin expect actual &&
    -+	rm expect actual
    ++	test_cmp_bin expect actual
      '
      
      test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message does not cause allocation failure' '
2:  0836f6b372 < -:  ---------- ci: match Linux jobs to available CPUs
-:  ---------- > 2:  8ec0308f6a ci: align job counts across CI providers

--- base-commit: 3bc0341126508f78f5869cbfc0005e987efdf0c7 change-id: 20260923-ci-large-test-resources-349cdc95f7f8

Tamir DubersteinSep 25, 2026, 16:35 UTC in reply to Tamir Duberstein on lore

[PATCH v2 1/2] t4205: compare huge output without diff

The huge-commit test compares output containing a line larger than 2 GiB. For two identical files containing 2,147,483,649 "1" bytes followed by "0\n", GNU diffutils 3.8 on Linux arm64 gives these measurements:

  Command               Mean +/- stddev       Maximum RSS (KiB)
  diff -u expect actual  5.276 +/- 0.572 s              4199924
  cmp expect actual      0.506 +/- 0.099 s                 1264

The test needs only an equality check. Use test_cmp_bin, which runs cmp, to compare the output byte for byte with less time and memory.

Assisted-by: LLM
Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
 t/t4205-log-pretty-formats.sh | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)
Show changes to t/t4205-log-pretty-formats.sh +1 −1
diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
index 4be5c51489..01b97c8888 100755
--- a/t/t4205-log-pretty-formats.sh
+++ b/t/t4205-log-pretty-formats.sh
@@ -1189,7 +1189,7 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
 test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
 	git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
 	echo 0 >>expect &&
-	test_cmp expect actual
+	test_cmp_bin expect actual
 '
 
 test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message does not cause allocation failure' '
-- 
2.56.0.rc2.815.g5b995412e4.frankengit
Tamir DubersteinSep 25, 2026, 16:35 UTC in reply to Tamir Duberstein on lore

[PATCH v2 2/2] ci: align job counts across CI providers

GitHub Actions sets JOBS to ten regardless of runner size, while GitLab CI uses the detected CPU count. Use the CPU count for Make and prove on both providers, selecting JOBS after the operating system is identified.

Use nproc on Linux and NUMBER_OF_PROCESSORS on Windows. On macOS, use sysctl to avoid requiring nproc before the dependency installer has run; GitHub macOS images need not provide GNU coreutils.

Assisted-by: LLM
Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
 ci/lib.sh | 16 ++++++++++++----
 1 file changed, 12 insertions(+), 4 deletions(-)
Show changes to ci/lib.sh +12 −4
diff --git a/ci/lib.sh b/ci/lib.sh
index c6ccbf8c17..db593cc62c 100755
--- a/ci/lib.sh
+++ b/ci/lib.sh
@@ -227,7 +227,6 @@ then
 	cache_dir="$HOME/none"
 
 	GIT_TEST_OPTS="--github-workflow-markup"
-	JOBS=10
 
 	distro=$(echo "$CI_JOB_IMAGE" | tr : -)
 elif test true = "$GITLAB_CI"
@@ -250,7 +249,6 @@ then
 	case "$OS,$CI_JOB_IMAGE" in
 	Windows_NT,*)
 		CI_OS_NAME=windows
-		JOBS=$NUMBER_OF_PROCESSORS
 		;;
 	*,macos-*)
 		# GitLab CI has Python installed via multiple package managers,
@@ -260,11 +258,9 @@ then
 		export PATH="$(brew --prefix)/bin:$PATH"
 
 		CI_OS_NAME=osx
-		JOBS=$(nproc)
 		;;
 	*,almalinux:*|*,alpine:*|*,debian:*|*,fedora:*|*,ubuntu:*|*,i386/ubuntu:*)
 		CI_OS_NAME=linux
-		JOBS=$(nproc)
 		;;
 	*)
 		echo "Could not identify OS image" >&2
@@ -291,6 +287,18 @@ else
 	exit 1
 fi
 
+case "$CI_OS_NAME" in
+windows|windows_nt)
+	JOBS=$NUMBER_OF_PROCESSORS
+	;;
+osx)
+	JOBS=$(sysctl -n hw.logicalcpu)
+	;;
+*)
+	JOBS=$(nproc)
+	;;
+esac
+
 MAKEFLAGS="$MAKEFLAGS --jobs=$JOBS"
 GIT_PROVE_OPTS="--timer --jobs $JOBS"
 
-- 
2.56.0.rc2.815.g5b995412e4.frankengit
Patrick SteinhardtSep 28, 2026, 06:36 UTC in reply to Tamir Duberstein on lore

Re: [PATCH 1/2] t4205: compare huge output without diff

On Thu, Sep 24, 2026 at 04:41:05PM -0400, Tamir Duberstein wrote:
Show 12 quoted lines
> On Thu, Sep 24, 2026 at 2:13 AM Patrick Steinhardt <ps@pks.im> wrote:
> >
> > On Wed, Sep 23, 2026 at 01:13:28PM -0400, Tamir Duberstein wrote:
> > > The huge-commit test compares two files with a line larger than 2 GiB.
> > > In Linux GitHub Actions jobs, git log produces its huge output but
> > > its subsequent diff process is killed with SIGKILL.
> >
> > I've never seen that failure before. Do you maybe have a link to it?
> 
> The failures happened on a private repo that I've since lost access to
> - but I believe it was precipitated by GitHub runners having half the
> memory in private repos as in public ones [1].
Okay.
Show 14 quoted lines
> > > Use test_cmp_bin to compare the output byte for byte without constructing
> > > a line-oriented diff. Remove the two large files after a successful
> > > comparison, releasing more than 4 GiB before subsequent tests.
> >
> > It would be great to back up the claim that test_cmp_bin is better than
> > test_cmp, e.g. by comparing peak RSS and its runtime.
> 
> As for the comparison: on Linux arm64 with GNU
> diffutils 3.8 using two identical files containing 2,147,483,649 "1" bytes
> followed by "0\n" (matching this test's expected output) gave:
> 
> Command              Mean +/- stddev       Maximum RSS (KiB)
> diff -u expect actual  5.276 +/- 0.572 s              4199924
> cmp expect actual      0.506 +/- 0.099 s                 1264
Quite a significant win indeed.
Show 25 quoted lines
> > > Signed-off-by: Tamir Duberstein <tamird@gmail.com>
> > > ---
> > >  t/t4205-log-pretty-formats.sh | 3 ++-
> > >  1 file changed, 2 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
> > > index 4be5c51489..6279a7e9bc 100755
> > > --- a/t/t4205-log-pretty-formats.sh
> > > +++ b/t/t4205-log-pretty-formats.sh
> > > @@ -1189,7 +1189,8 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
> > >  test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
> > >       git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
> > >       echo 0 >>expect &&
> > > -     test_cmp expect actual
> > > +     test_cmp_bin expect actual &&
> > > +     rm expect actual
> > >  '
> >
> > Hm. Sure, releasing these files isn't a bad idea by itself. But we
> > rewrite "expect" in the next test anyway, and "actual" will be rewritten
> > two tests further down. So does it really buy us that much...?
> 
> You're right, this probably does not buy much.
> 
> Would you like me to include the performance comparison in v2?

I think that'd be good, yes. Providing context like this to the reviewer makes everyone's life easier :)

> As for the deletion: would you prefer I drop it?

My personal take is that we can just drop it as it doesn't buy us much. If we want to keep it we should be honest about its effect in the commit message.

Patrick
Patrick SteinhardtSep 28, 2026, 06:36 UTC in reply to Tamir Duberstein on lore

Re: [PATCH v2 1/2] t4205: compare huge output without diff

On Fri, Sep 25, 2026 at 12:35:38PM -0400, Tamir Duberstein wrote:
Show 10 quoted lines
> The huge-commit test compares output containing a line larger than 2 GiB.
> For two identical files containing 2,147,483,649 "1" bytes followed by
> "0\n", GNU diffutils 3.8 on Linux arm64 gives these measurements:
> 
>   Command               Mean +/- stddev       Maximum RSS (KiB)
>   diff -u expect actual  5.276 +/- 0.572 s              4199924
>   cmp expect actual      0.506 +/- 0.099 s                 1264
> 
> The test needs only an equality check. Use test_cmp_bin, which runs cmp,
> to compare the output byte for byte with less time and memory.
Yup, this is much more compelling as an argument now :)
Show 11 quoted lines
> diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
> index 4be5c51489..01b97c8888 100755
> --- a/t/t4205-log-pretty-formats.sh
> +++ b/t/t4205-log-pretty-formats.sh
> @@ -1189,7 +1189,7 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
>  test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
>  	git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
>  	echo 0 >>expect &&
> -	test_cmp expect actual
> +	test_cmp_bin expect actual
>  '
And the patch looks obviously good to me.
Patrick
Patrick SteinhardtSep 28, 2026, 06:36 UTC in reply to Tamir Duberstein on lore

Re: [PATCH v2 2/2] ci: align job counts across CI providers

On Fri, Sep 25, 2026 at 12:35:39PM -0400, Tamir Duberstein wrote:
> GitHub Actions sets JOBS to ten regardless of runner size, while
> GitLab CI uses the detected CPU count. Use the CPU count for Make and
> prove on both providers, selecting JOBS after the operating system
> is identified.

Again, it should be noted here what the effect of this is. In other words, does GitHub slow down as a result? You already showed numbers during the discussion on v1 of this series, and these numbers should probably be included in this message, too.

> Use nproc on Linux and NUMBER_OF_PROCESSORS on Windows. On macOS, use
> sysctl to avoid requiring nproc before the dependency installer has run;
> GitHub macOS images need not provide GNU coreutils.

Huh... "need not" feels somewhat weird as phrasing. I guess it's rather "does not", and consequently we have to adapt? I think instead of describing what you do, I'd directly pinpoint what matters:

  Note that we continue to use the same logic to detect the number of
  processors on both Linux and Windows. But on macOS, we cannot continue
  to use nproc(1) because the image used by GitHub does not provide that
  tool. Use sysctl instead, which is available on both GitLab and
  GitHub.
Thanks!
Patrick
Tamir DubersteinSep 28, 2026, 10:16 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH v2 2/2] ci: align job counts across CI providers

On Mon, Sep 28, 2026 at 2:36 AM Patrick Steinhardt <ps@pks.im> wrote:
Show 11 quoted lines
>
> On Fri, Sep 25, 2026 at 12:35:39PM -0400, Tamir Duberstein wrote:
> > GitHub Actions sets JOBS to ten regardless of runner size, while
> > GitLab CI uses the detected CPU count. Use the CPU count for Make and
> > prove on both providers, selecting JOBS after the operating system
> > is identified.
>
> Again, it should be noted here what the effect of this is. In other
> words, does GitHub slow down as a result? You already showed numbers
> during the discussion on v1 of this series, and these numbers should
> probably be included in this message, too.

Agreed, but in this case there was no reliable performance change across 10 runs; I could include that.

Show 14 quoted lines
>
> > Use nproc on Linux and NUMBER_OF_PROCESSORS on Windows. On macOS, use
> > sysctl to avoid requiring nproc before the dependency installer has run;
> > GitHub macOS images need not provide GNU coreutils.
>
> Huh... "need not" feels somewhat weird as phrasing. I guess it's rather
> "does not", and consequently we have to adapt? I think instead of
> describing what you do, I'd directly pinpoint what matters:
>
>   Note that we continue to use the same logic to detect the number of
>   processors on both Linux and Windows. But on macOS, we cannot continue
>   to use nproc(1) because the image used by GitHub does not provide that
>   tool. Use sysctl instead, which is available on both GitLab and
>   GitHub

Agreed. .

>
> Thanks!
>
> Patrick
Patrick SteinhardtSep 28, 2026, 11:14 UTC in reply to Tamir Duberstein on lore

Re: [PATCH v2 2/2] ci: align job counts across CI providers

On Mon, Sep 28, 2026 at 06:16:33AM -0400, Tamir Duberstein wrote:
Show 15 quoted lines
> On Mon, Sep 28, 2026 at 2:36 AM Patrick Steinhardt <ps@pks.im> wrote:
> >
> > On Fri, Sep 25, 2026 at 12:35:39PM -0400, Tamir Duberstein wrote:
> > > GitHub Actions sets JOBS to ten regardless of runner size, while
> > > GitLab CI uses the detected CPU count. Use the CPU count for Make and
> > > prove on both providers, selecting JOBS after the operating system
> > > is identified.
> >
> > Again, it should be noted here what the effect of this is. In other
> > words, does GitHub slow down as a result? You already showed numbers
> > during the discussion on v1 of this series, and these numbers should
> > probably be included in this message, too.
> 
> Agreed, but in this case there was no reliable performance change
> across 10 runs; I could include that.

I think it should be included, as it's the one thing that people will be wondering about when they see this change.

Patrick
Tamir DubersteinSep 30, 2026, 14:20 UTC in reply to Tamir Duberstein on lore

[PATCH v3 0/2] ci: use cmp and align job-count selection

The first patch uses cmp for the huge commit-message output comparison. On this input, GNU diffutils 3.8 on Linux arm64 takes 5.276 seconds with 4,199,924 KiB peak RSS for diff, versus 0.506 seconds and 1,264 KiB for cmp. Patch 1 includes the fixture and measurement details.

The second patch uses twice the detected CPU count for Make and prove on both GitHub Actions and GitLab CI. This replaces GitHub's fixed ten jobs and doubles GitLab's job count. The GitHub measurements in patch 2 favor 2*N over N, with mixed results against ten jobs. GitLab runtime and resource use have not been measured.

Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
The table in patch 2 uses attempts 1-5 of each linked run. Its rows
aggregate ten Linux jobs, three macOS jobs, and a Windows build plus ten
test shards. Failed steps are excluded. Each job has five samples,
except one CPU-count Windows shard and one 2x-CPU Linux job with four.

Five further attempts per policy for osx-clang on three-CPU macOS runners gave these successful build/test-step medians (minutes:seconds):

  Jobs                3       6       10
  Median          48:10   36:44    38:05.5
  Passed              5       5        4
  Cancelled           0       0        1

Six jobs had lower times than three in each block. The ten-job cancellation followed six hours in the build/test step; its cause is unknown. It is excluded from the successful-duration median above. These are attempts 6-10 of the runs linked in patch 2, separate from its earlier full CI matrix measurements. Neither experiment measured peak memory or disk use.

Changes in v3:
- Use twice the CPU count on both providers, instead of adopting
  GitLab's existing one-job-per-CPU policy.
- Include CI timings and their tradeoffs in patch 2's commit message,
  and explain the use of sysctl directly.
- Patch 1 is unchanged.
- Link to v2: https://patch.msgid.link/20260925-ci-large-test-resources-v2-0-f632cf319756@gmail.com
Changes in v2:
- Replace the unavailable CI failure reference with comparison runtime
  and peak RSS measurements.
- Drop the file removal; following tests overwrite expect and actual.
- Share job-count selection between GitHub Actions and GitLab CI.
- Use native CPU-count queries on macOS and Windows.
- Link to v1: https://patch.msgid.link/20260923-ci-large-test-resources-v1-0-c28416d59475@gmail.com
---
Tamir Duberstein (2):
      t4205: compare huge output without diff
      ci: use twice the CPU count on both providers
 ci/lib.sh                     | 17 +++++++++++++----
 t/t4205-log-pretty-formats.sh |  2 +-
 2 files changed, 14 insertions(+), 5 deletions(-)

--- Range-diff versus v2:

1:  5cf348a75c = 1:  4a15b8f17d t4205: compare huge output without diff
2:  3b9bd2c495 ! 2:  d444411070 ci: align job counts across CI providers
    @@ Metadata
     Author: Tamir Duberstein <tamird@gmail.com>
     
      ## Commit message ##
    -    ci: align job counts across CI providers
    +    ci: use twice the CPU count on both providers
     
         GitHub Actions sets JOBS to ten regardless of runner size, while
    -    GitLab CI uses the detected CPU count. Use the CPU count for Make and
    -    prove on both providers, selecting JOBS after the operating system
    -    is identified.
    +    GitLab CI uses the detected CPU count. Use twice the CPU count for
    +    Make and prove on both providers, doubling GitLab's job count.
     
    -    Use nproc on Linux and NUMBER_OF_PROCESSORS on Windows. On macOS, use
    -    sysctl to avoid requiring nproc before the dependency installer has run;
    -    GitHub macOS images need not provide GNU coreutils.
    +    Five GitHub Actions attempts per policy, with the long tests enabled,
    +    gave these sums of per-job median successful build/test-step times
    +    (minutes; four or five samples per job) [1-3]:
    +
    +                          Fixed 10   CPU count   2x CPU count
    +      Linux Make             278.9       273.9          259.9
    +      macOS Make              94.5       119.1           99.8
    +      Windows Make           102.2       103.8          100.8
    +
    +    Workflow overhead is excluded; Windows runner images varied.
    +
    +    Use twice the CPU count to scale concurrency with runner size while
    +    avoiding the larger macOS slowdown observed with one job per CPU.
    +    Compared with ten jobs, this trades a lower Linux total for a higher
    +    macOS total.
    +
    +    Keep GitLab's Linux and Windows CPU queries. On macOS, use the native
    +    sysctl command on both providers so CPU detection does not depend on
    +    GNU coreutils.
    +
    +    Link: https://github.com/tamird/git/actions/runs/36070869894/attempts/1 [1]
    +    Link: https://github.com/tamird/git/actions/runs/36070867524/attempts/1 [2]
    +    Link: https://github.com/tamird/git/actions/runs/36070867647/attempts/1 [3]
     
         Assisted-by: LLM
         Signed-off-by: Tamir Duberstein <tamird@gmail.com>
    @@ ci/lib.sh: else
     +	JOBS=$(nproc)
     +	;;
     +esac
    ++JOBS=$((2 * JOBS))
     +
      MAKEFLAGS="$MAKEFLAGS --jobs=$JOBS"
      GIT_PROVE_OPTS="--timer --jobs $JOBS"

--- base-commit: 3bc0341126508f78f5869cbfc0005e987efdf0c7 change-id: 20260923-ci-large-test-resources-349cdc95f7f8

Tamir DubersteinSep 30, 2026, 14:20 UTC in reply to Tamir Duberstein on lore

[PATCH v3 1/2] t4205: compare huge output without diff

The huge-commit test compares output containing a line larger than 2 GiB. For two identical files containing 2,147,483,649 "1" bytes followed by "0\n", GNU diffutils 3.8 on Linux arm64 gives these measurements:

  Command               Mean +/- stddev       Maximum RSS (KiB)
  diff -u expect actual  5.276 +/- 0.572 s              4199924
  cmp expect actual      0.506 +/- 0.099 s                 1264

The test needs only an equality check. Use test_cmp_bin, which runs cmp, to compare the output byte for byte with less time and memory.

Assisted-by: LLM
Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
 t/t4205-log-pretty-formats.sh | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)
Show changes to t/t4205-log-pretty-formats.sh +1 −1
diff --git a/t/t4205-log-pretty-formats.sh b/t/t4205-log-pretty-formats.sh
index 4be5c51489..01b97c8888 100755
--- a/t/t4205-log-pretty-formats.sh
+++ b/t/t4205-log-pretty-formats.sh
@@ -1189,7 +1189,7 @@ test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'set up huge commit' '
 test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message' '
 	git log -1 --format="%B%<(1)%x30" $huge_commit >actual &&
 	echo 0 >>expect &&
-	test_cmp expect actual
+	test_cmp_bin expect actual
 '
 
 test_expect_success EXPENSIVE,SIZE_T_IS_64BIT 'log --pretty with huge commit message does not cause allocation failure' '
-- 
2.56.0.rc2.851.g50a151a6e9.frankengit
Tamir DubersteinSep 30, 2026, 14:20 UTC in reply to Tamir Duberstein on lore

[PATCH v3 2/2] ci: use twice the CPU count on both providers

GitHub Actions sets JOBS to ten regardless of runner size, while GitLab CI uses the detected CPU count. Use twice the CPU count for Make and prove on both providers, doubling GitLab's job count.

Five GitHub Actions attempts per policy, with the long tests enabled, gave these sums of per-job median successful build/test-step times (minutes; four or five samples per job) [1-3]:

                      Fixed 10   CPU count   2x CPU count
  Linux Make             278.9       273.9          259.9
  macOS Make              94.5       119.1           99.8
  Windows Make           102.2       103.8          100.8
Workflow overhead is excluded; Windows runner images varied.

Use twice the CPU count to scale concurrency with runner size while avoiding the larger macOS slowdown observed with one job per CPU. Compared with ten jobs, this trades a lower Linux total for a higher macOS total.

Keep GitLab's Linux and Windows CPU queries. On macOS, use the native sysctl command on both providers so CPU detection does not depend on GNU coreutils.

Link: https://github.com/tamird/git/actions/runs/36070869894/attempts/1 [1]
Link: https://github.com/tamird/git/actions/runs/36070867524/attempts/1 [2]
Link: https://github.com/tamird/git/actions/runs/36070867647/attempts/1 [3]
Assisted-by: LLM
Signed-off-by: Tamir Duberstein <tamird@gmail.com>
---
 ci/lib.sh | 17 +++++++++++++----
 1 file changed, 13 insertions(+), 4 deletions(-)
Show changes to ci/lib.sh +13 −4
diff --git a/ci/lib.sh b/ci/lib.sh
index c6ccbf8c17..f0b7f35850 100755
--- a/ci/lib.sh
+++ b/ci/lib.sh
@@ -227,7 +227,6 @@ then
 	cache_dir="$HOME/none"
 
 	GIT_TEST_OPTS="--github-workflow-markup"
-	JOBS=10
 
 	distro=$(echo "$CI_JOB_IMAGE" | tr : -)
 elif test true = "$GITLAB_CI"
@@ -250,7 +249,6 @@ then
 	case "$OS,$CI_JOB_IMAGE" in
 	Windows_NT,*)
 		CI_OS_NAME=windows
-		JOBS=$NUMBER_OF_PROCESSORS
 		;;
 	*,macos-*)
 		# GitLab CI has Python installed via multiple package managers,
@@ -260,11 +258,9 @@ then
 		export PATH="$(brew --prefix)/bin:$PATH"
 
 		CI_OS_NAME=osx
-		JOBS=$(nproc)
 		;;
 	*,almalinux:*|*,alpine:*|*,debian:*|*,fedora:*|*,ubuntu:*|*,i386/ubuntu:*)
 		CI_OS_NAME=linux
-		JOBS=$(nproc)
 		;;
 	*)
 		echo "Could not identify OS image" >&2
@@ -291,6 +287,19 @@ else
 	exit 1
 fi
 
+case "$CI_OS_NAME" in
+windows|windows_nt)
+	JOBS=$NUMBER_OF_PROCESSORS
+	;;
+osx)
+	JOBS=$(sysctl -n hw.logicalcpu)
+	;;
+*)
+	JOBS=$(nproc)
+	;;
+esac
+JOBS=$((2 * JOBS))
+
 MAKEFLAGS="$MAKEFLAGS --jobs=$JOBS"
 GIT_PROVE_OPTS="--timer --jobs $JOBS"
 
-- 
2.56.0.rc2.851.g50a151a6e9.frankengit
Patrick SteinhardtSep 30, 2026, 14:45 UTC in reply to Tamir Duberstein on lore

Re: [PATCH v3 2/2] ci: use twice the CPU count on both providers

On Wed, Sep 30, 2026 at 10:20:35AM -0400, Tamir Duberstein wrote:
Show 19 quoted lines
> GitHub Actions sets JOBS to ten regardless of runner size, while
> GitLab CI uses the detected CPU count. Use twice the CPU count for
> Make and prove on both providers, doubling GitLab's job count.
> 
> Five GitHub Actions attempts per policy, with the long tests enabled,
> gave these sums of per-job median successful build/test-step times
> (minutes; four or five samples per job) [1-3]:
> 
>                       Fixed 10   CPU count   2x CPU count
>   Linux Make             278.9       273.9          259.9
>   macOS Make              94.5       119.1           99.8
>   Windows Make           102.2       103.8          100.8
> 
> Workflow overhead is excluded; Windows runner images varied.
> 
> Use twice the CPU count to scale concurrency with runner size while
> avoiding the larger macOS slowdown observed with one job per CPU.
> Compared with ten jobs, this trades a lower Linux total for a higher
> macOS total.

Right. We could of course special-case macOS. But I don't feel like it makes sense to squeeze every single second out of a job that's already the fastest anyway.

Patrick
Patrick SteinhardtSep 30, 2026, 14:45 UTC in reply to Tamir Duberstein on lore

Re: [PATCH v3 0/2] ci: use cmp and align job-count selection

On Wed, Sep 30, 2026 at 10:20:33AM -0400, Tamir Duberstein wrote:
Show 7 quoted lines
> Changes in v3:
> - Use twice the CPU count on both providers, instead of adopting
>   GitLab's existing one-job-per-CPU policy.
> - Include CI timings and their tradeoffs in patch 2's commit message,
>   and explain the use of sysctl directly.
> - Patch 1 is unchanged.
> - Link to v2: https://patch.msgid.link/20260925-ci-large-test-resources-v2-0-f632cf319756@gmail.com
Thanks, I'm happy with this version.
Patrick
Tamir DubersteinSep 30, 2026, 14:58 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH v3 0/2] ci: use cmp and align job-count selection

On Wed, Sep 30, 2026 at 10:45 AM Patrick Steinhardt <ps@pks.im> wrote:
Show 13 quoted lines
>
> On Wed, Sep 30, 2026 at 10:20:33AM -0400, Tamir Duberstein wrote:
> > Changes in v3:
> > - Use twice the CPU count on both providers, instead of adopting
> >   GitLab's existing one-job-per-CPU policy.
> > - Include CI timings and their tradeoffs in patch 2's commit message,
> >   and explain the use of sysctl directly.
> > - Patch 1 is unchanged.
> > - Link to v2: https://patch.msgid.link/20260925-ci-large-test-resources-v2-0-f632cf319756@gmail.com
>
> Thanks, I'm happy with this version.
>
> Patrick
Thanks for the reviews!
Junio C HamanoSep 30, 2026, 18:04 UTC in reply to Tamir Duberstein on lore

Re: [PATCH v3 0/2] ci: use cmp and align job-count selection

Tamir Duberstein <tamird@gmail.com> writes:
Show 16 quoted lines
> On Wed, Sep 30, 2026 at 10:45 AM Patrick Steinhardt <ps@pks.im> wrote:
>>
>> On Wed, Sep 30, 2026 at 10:20:33AM -0400, Tamir Duberstein wrote:
>> > Changes in v3:
>> > - Use twice the CPU count on both providers, instead of adopting
>> >   GitLab's existing one-job-per-CPU policy.
>> > - Include CI timings and their tradeoffs in patch 2's commit message,
>> >   and explain the use of sysctl directly.
>> > - Patch 1 is unchanged.
>> > - Link to v2: https://patch.msgid.link/20260925-ci-large-test-resources-v2-0-f632cf319756@gmail.com
>>
>> Thanks, I'm happy with this version.
>>
>> Patrick
>
> Thanks for the reviews!
Thanks for writing and reviewing these patches, both of you.
Let me mark it for 'next'.

Back to recent threads