Re: [PATCH v3] sparse-checkout: optimize string_list construction
- From
Derrick Stolee <stolee@gmail.com>
- Date
- Jan 18, 2026, 02:46 UTC
- Message-ID
- <c5631f7d-72ff-4876-9b68-ea4a70fde501@gmail.com>
- In-Reply-To
- <xmqqy0lx8ojt.fsf@gitster.g>
On 1/16/26 12:17 PM, Junio C Hamano wrote:
Show 13 quoted lines
> Amisha Chhajed <amishhhaaaa@gmail.com> writes: > >> It was assumed to be safe under the notion that our entries are not >> duplicate but as already pointed out, our entries are not unique so we >> need one of those two ways either insert or remove_duplicates, this >> can be a trivial question but i wonder how are the tests passing by >> removing these lines, i was actually researching about it. > > ... suspense. And the result of the research was??? > > If the answer was simply "we lack test coverage", it may make sense > to add a test taken from Peff's earlier response to increase test > coverage, perhaps?
In addition to adding more tests to t/t1091-sparse-checkout-builtin.sh to cover these duplicate cases. To demonstrate your quadratic perf improvement, a test in t/perf/p2000-sparse-operations.sh or similar would be good to add.
I expect that the test you would add doesn't matter too much about the data shape, but would look very different from most tests in p2000. You can make use of the constructed repo's directory structure that has nesting directories with name f1, f2, f3, or f4.
Here's something to get you started that I haven't tested myself:
test_perf 'duplicate sparse directories' ' ( cd full-v4 && for i in $(test_seq 1000) do printf "f1/f2/f3/f4\n" done >in && git sparse-checkout set --stdin <in ) '
That should test the logic with 1000 identical directories, which should be enough to have the quadratic growth show up.
Thanks, -Stolee