From: Amisha Chhajed Date: Sun, 18 Jan 2026 13:09:58 GMT Subject: Re: [PATCH v3] sparse-checkout: optimize string_list construction Message-ID: In-Reply-To: On Sun, 18 Jan 2026 at 08:16, Derrick Stolee wrote: > > On 1/16/26 12:17 PM, Junio C Hamano wrote: > > Amisha Chhajed writes: > > > >> It was assumed to be safe under the notion that our entries are not > >> duplicate but as already pointed out, our entries are not unique so we > >> need one of those two ways either insert or remove_duplicates, this > >> can be a trivial question but i wonder how are the tests passing by > >> removing these lines, i was actually researching about it. > > > > ... suspense. And the result of the research was??? > > > > If the answer was simply "we lack test coverage", it may make sense > > to add a test taken from Peff's earlier response to increase test > > coverage, perhaps? > > In addition to adding more tests to t/t1091-sparse-checkout-builtin.sh > to cover these duplicate cases. To demonstrate your quadratic perf > improvement, a test in t/perf/p2000-sparse-operations.sh or similar > would be good to add. > > I expect that the test you would add doesn't matter too much about > the data shape, but would look very different from most tests in > p2000. You can make use of the constructed repo's directory structure > that has nesting directories with name f1, f2, f3, or f4. > > Here's something to get you started that I haven't tested myself: > > test_perf 'duplicate sparse directories' ' > ( > cd full-v4 && > > for i in $(test_seq 1000) > do > printf "f1/f2/f3/f4\n" > done >in && > git sparse-checkout set --stdin ) > ' > > That should test the logic with 1000 identical directories, which > should be enough to have the quadratic growth show up. > > Thanks, > -Stolee Thank you so much for pointing me in the right direction.