git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: t7900's new expensive test

From
Derrick Stolee <stolee@gmail.com>
Date
Dec 1, 2020, 20:55 UTC
Message-ID
<373f3dfe-828b-430d-b88e-5e23302090cb@gmail.com>
In-Reply-To
<X8YrbDpC9/EjRr95@coredump.intra.peff.net>
On 12/1/2020 6:39 AM, Jeff King wrote:
Show 11 quoted lines
> On Tue, Dec 01, 2020 at 06:23:28AM -0500, Jeff King wrote:
> 
>> I'm not sure if EXPENSIVE is the right ballpark, or if we'd want a
>> VERY_EXPENSIVE. On my machine, the whole test suite for v2.29.0 takes 64
>> seconds to run, and setting GIT_TEST_LONG=1 bumps that to 103s. It got a
>> bit worse since then, as t7900 adds an EXPENSIVE test that takes ~200s
>> (it's not strictly additive, since we can work in parallel on other
>> tests for the first bit, but still, yuck).
> 
> Since Stolee is on the cc and has already seen me complaining about his
> test, I guess I should expand a bit. ;)

Ha. I apologize for causing pain here. My thought was that GIT_TEST_LONG=1 was only used by someone really willing to wait, or someone specifically trying to investigate a problem that only triggers on very large cases.

In that sense, it's not so much intended as a frequently-run regression test, but a "run this if you are messing with this area" kind of thing. Perhaps there is a different pattern to use here?

Show 12 quoted lines
> There are some small wins possible (e.g., using "commit --quiet" seems
> to shave off ~8s when we don't even think about writing a diff), but
> fundamentally the issue is that it just takes a long time to "git add"
> the 5.2GB worth of random data. I almost wonder if it would be worth it
> to hard-coded the known sha1 and sha256 names of the blobs, and write
> them straight into the appropriate loose object file. I guess that is
> tricky, though, because it actually needs to be a zlib stream, not just
> the output of "test-tool genrandom".
>
> Though speaking of which, another easy win might be setting
> core.compression to "0". We know the random data won't compress anyway,
> so there's no point in spending cycles on zlib.
The intention is mostly to expand the data beyond two gigabytes, so
dropping compression to get there seems like a good idea. If we are
not compressing at all, then perhaps we can reliably cut ourselves
closer to the 2GB limit instead of overshooting as a precaution.
 
Show 29 quoted lines
> Doing this:
> 
> diff --git a/t/t7900-maintenance.sh b/t/t7900-maintenance.sh
> index d9e68bb2bf..849c6d1361 100755
> --- a/t/t7900-maintenance.sh
> +++ b/t/t7900-maintenance.sh
> @@ -239,6 +239,8 @@ test_expect_success 'incremental-repack task' '
>  '
>  
>  test_expect_success EXPENSIVE 'incremental-repack 2g limit' '
> +	test_config core.compression 0 &&
> +
>  	for i in $(test_seq 1 5)
>  	do
>  		test-tool genrandom foo$i $((512 * 1024 * 1024 + 1)) >>big ||
> @@ -257,7 +259,7 @@ test_expect_success EXPENSIVE 'incremental-repack 2g limit' '
>  		return 1
>  	done &&
>  	git add big &&
> -	git commit -m "Add big file (2)" &&
> +	git commit -qm "Add big file (2)" &&
>  
>  	# ensure any possible loose objects are in a pack-file
>  	git maintenance run --task=loose-objects &&
> 
> seems to shave off ~140s from the test. I think we could get a little
> more by cleaning up the enormous objects, too (they end up causing the
> subsequent test to run slower, too, though perhaps it was intentional to
> impact downstream tests).

Cutting out 70% out seems like a great idea. I don't think it was super intentional to slow down those tests.

Thanks, -Stolee

Previous: Jeff KingNext: Jeff King
Message 15 of 19 in “handling 4GB .idx files”
  1. 0/5 handling 4GB .idx filesJeff King, Nov 13, 2020
  2. 1/5 compute pack .idx byte offsets using size_tJeff King, Nov 13, 2020
  3. 2/5 use size_t to store pack .idx byte offsetsJeff King, Nov 13, 2020
  4. 3/5 fsck: correctly compute checksums on idx files larger than 4GBJeff King, Nov 13, 2020
  5. 4/5 block-sha1: take a size_t length parameterJeff King, Nov 13, 2020
  6. 5/5 packfile: detect overflow in .idx file size checksJeff King, Nov 13, 2020
  7. Johannes SchindelinNov 13, 2020
  8. Thomas BraunNov 15, 2020
  9. Jeff KingNov 16, 2020
  10. Derrick StoleeNov 16, 2020
  11. Jeff KingNov 16, 2020
  12. Thomas BraunNov 30, 2020
  13. Jeff KingDec 1, 2020
  14. t7900's new expensive testJeff King, Dec 1, 2020
  15. Derrick StoleeDec 1, 2020
  16. t7900: speed up expensive testJeff King, Dec 2, 2020
  17. Derrick StoleeDec 3, 2020
  18. Taylor BlauDec 1, 2020
  19. Jeff KingDec 2, 2020

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.