# [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD

6 messages from 2026-08-02 to 2026-09-11. Participants: Sahitya Chandra, Elijah Newren.
Thread: https://gitlist.dev/t/66102

## Sahitya Chandra, 2026-08-02 21:28

Subject: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD
Message-ID: <20260802212826.1090943-1-sahityajb@gmail.com>

```
repo_index_has_changes() normally checks whether the index differs from
a tree by passing that tree to the diff machinery. When no tree is
passed, it tries to use HEAD for that comparison.

If HEAD does not resolve, as on an unborn branch, the function falls
back to walking the index directly. With a sparse index, however, sparse
directory entries may stand in for many paths, so the fallback first
expands the index before reporting the changed paths.

That expansion is unnecessary. An unborn HEAD is equivalent for this
check to comparing the index against the empty tree: every index entry
is new relative to that tree.

Use the empty tree when HEAD cannot be resolved. This keeps the
unborn-branch case on the same diff code path as the normal
tree-comparison case, avoiding the sparse-index expansion while still
letting callers see paths inside sparse directories.

Teach test-tool read-cache to exercise repo_index_has_changes(), and
add a t1092 check that the unborn-branch case reports paths inside a
sparse directory without expanding the index.

Signed-off-by: Sahitya Chandra <sahityajb@gmail.com>
---
 read-cache.c                             | 46 ++++++++++--------------
 t/helper/test-read-cache.c               | 19 ++++++++++
 t/t1092-sparse-checkout-compatibility.sh | 16 +++++++++
 3 files changed, 53 insertions(+), 28 deletions(-)

diff --git a/read-cache.c b/read-cache.c
index 6c449f393d..88ee9ba935 100644
--- a/read-cache.c
+++ b/read-cache.c
@@ -2505,39 +2505,29 @@ int repo_index_has_changes(struct repository *repo,
 			   struct tree *tree,
 			   struct strbuf *sb)
 {
-	struct index_state *istate = repo->index;
+	struct diff_options opt;
 	struct object_id cmp;
 	int i;
 
 	if (tree)
 		cmp = tree->object.oid;
-	if (tree || !repo_get_oid_tree(repo, "HEAD", &cmp)) {
-		struct diff_options opt;
-
-		repo_diff_setup(repo, &opt);
-		opt.flags.exit_with_status = 1;
-		if (!sb)
-			opt.flags.quick = 1;
-		diff_setup_done(&opt);
-		do_diff_cache(&cmp, &opt);
-		diffcore_std(&opt);
-		for (i = 0; sb && i < diff_queued_diff.nr; i++) {
-			if (i)
-				strbuf_addch(sb, ' ');
-			strbuf_addstr(sb, diff_queued_diff.queue[i]->two->path);
-		}
-		diff_flush(&opt);
-		return opt.flags.has_changes != 0;
-	} else {
-		/* TODO: audit for interaction with sparse-index. */
-		ensure_full_index(istate);
-		for (i = 0; sb && i < istate->cache_nr; i++) {
-			if (i)
-				strbuf_addch(sb, ' ');
-			strbuf_addstr(sb, istate->cache[i]->name);
-		}
-		return !!istate->cache_nr;
-	}
+	else if (repo_get_oid_tree(repo, "HEAD", &cmp))
+		oidcpy(&cmp, repo->hash_algo->empty_tree);
+
+	repo_diff_setup(repo, &opt);
+	opt.flags.exit_with_status = 1;
+	if (!sb)
+		opt.flags.quick = 1;
+	diff_setup_done(&opt);
+	do_diff_cache(&cmp, &opt);
+	diffcore_std(&opt);
+	for (i = 0; sb && i < diff_queued_diff.nr; i++) {
+		if (i)
+			strbuf_addch(sb, ' ');
+		strbuf_addstr(sb, diff_queued_diff.queue[i]->two->path);
+	}
+	diff_flush(&opt);
+	return opt.flags.has_changes != 0;
 }
 
 static int write_index_ext_header(struct hashfile *f,
diff --git a/t/helper/test-read-cache.c b/t/helper/test-read-cache.c
index 6b08ba8f07..ee629fbc69 100644
--- a/t/helper/test-read-cache.c
+++ b/t/helper/test-read-cache.c
@@ -4,6 +4,7 @@
 #include "config.h"
 #include "environment.h"
 #include "read-cache-ll.h"
+#include "repo-settings.h"
 #include "repository.h"
 #include "setup.h"
 
@@ -12,6 +13,24 @@ int cmd__read_cache(int argc, const char **argv)
 	int i, cnt = 1;
 	const char *name = NULL;
 
+	if (argc == 2 && !strcmp(argv[1], "--index-has-changes")) {
+		struct strbuf sb = STRBUF_INIT;
+		int ret;
+
+		setup_git_directory(the_repository);
+		repo_config(the_repository, git_default_config, NULL);
+		prepare_repo_settings(the_repository);
+		the_repository->settings.command_requires_full_index = 0;
+
+		repo_read_index(the_repository);
+		ret = repo_index_has_changes(the_repository, NULL, &sb);
+		printf("has_changes=%d\n", ret);
+		if (sb.len)
+			printf("dirty=%s\n", sb.buf);
+		strbuf_release(&sb);
+		return 0;
+	}
+
 	if (argc > 1 && skip_prefix(argv[1], "--print-and-refresh=", &name)) {
 		argc--;
 		argv++;
diff --git a/t/t1092-sparse-checkout-compatibility.sh b/t/t1092-sparse-checkout-compatibility.sh
index 4140c4d8ef..90239a862d 100755
--- a/t/t1092-sparse-checkout-compatibility.sh
+++ b/t/t1092-sparse-checkout-compatibility.sh
@@ -1558,6 +1558,22 @@ test_expect_success 'sparse-index is not expanded' '
 	)
 '
 
+test_expect_success 'sparse-index is not expanded: index has changes on unborn branch' '
+	init_repos &&
+	git -C sparse-index checkout --orphan unborn &&
+	git -C sparse-index ls-files --sparse --stage >cache &&
+	test_grep "^040000 .*	folder1/$" cache &&
+
+	rm -f trace2.txt &&
+	GIT_TRACE2_EVENT="$(pwd)/trace2.txt" GIT_TRACE2_EVENT_NESTING=10 \
+		test-tool -C sparse-index read-cache --index-has-changes \
+		>sparse-index-out 2>sparse-index-error &&
+	test_region ! index ensure_full_index trace2.txt &&
+	test_must_be_empty sparse-index-error &&
+	test_grep "has_changes=1" sparse-index-out &&
+	test_grep "folder1/a" sparse-index-out
+'
+
 test_expect_success 'sparse-index is not expanded: merge conflict in cone' '
 	init_repos &&
 

base-commit: a97fcc37c2bc6340a8d7ce78dedf227aac4e9aa7
-- 
2.43.0


```

## Sahitya Chandra, 2026-08-06 15:08

Subject: Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD
Message-ID: <CAP=WS+vys5ob20mkxpzPqUjeCqG6hm7-EeDdec0Y0NaBc+tT1A@mail.gmail.com>
In-Reply-To: <20260802212826.1090943-1-sahityajb@gmail.com>

```
Just a gentle ping on this patch.

Thanks,
Sahitya

```

## Elijah Newren, 2026-08-07 06:47

Subject: Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD
Message-ID: <CABPp-BGYuQA_ngR3xS-_Mndzf_ubkn7rSc25CJG=UbLCVGdnyg@mail.gmail.com>
In-Reply-To: <20260802212826.1090943-1-sahityajb@gmail.com>

```
On Sun, Aug 2, 2026 at 2:28 PM Sahitya Chandra <sahityajb@gmail.com> wrote:
>
> repo_index_has_changes() normally checks whether the index differs from
> a tree by passing that tree to the diff machinery. When no tree is
> passed, it tries to use HEAD for that comparison.
>
> If HEAD does not resolve, as on an unborn branch, the function falls
> back to walking the index directly. With a sparse index, however, sparse
> directory entries may stand in for many paths, so the fallback first
> expands the index before reporting the changed paths.
>
> That expansion is unnecessary. An unborn HEAD is equivalent for this
> check to comparing the index against the empty tree: every index entry
> is new relative to that tree.
>
> Use the empty tree when HEAD cannot be resolved. This keeps the
> unborn-branch case on the same diff code path as the normal
> tree-comparison case, avoiding the sparse-index expansion while still
> letting callers see paths inside sparse directories.
>
> Teach test-tool read-cache to exercise repo_index_has_changes(), and
> add a t1092 check that the unborn-branch case reports paths inside a
> sparse directory without expanding the index.

This explains what, but not why.  It feels like a pedagogical exercise
with no actual utility.  Why would someone with an unborn HEAD be
using a sparse index?  They have millions of files, with none of them
committed, except they don't have millions of files because they only
have paths under certain directories?  How did they even get the
relevant tree entries into the sparse index in order to have one?

Perhaps you have a great usecase and I've just missed it.  Could you
explain the motivation for enabling this?  Or was it more a case of
trying to take care of TODOs in the code?

[...]
> @@ -12,6 +13,24 @@ int cmd__read_cache(int argc, const char **argv)
>         int i, cnt = 1;
>         const char *name = NULL;
>
> +       if (argc == 2 && !strcmp(argv[1], "--index-has-changes")) {
> +               struct strbuf sb = STRBUF_INIT;
> +               int ret;
> +
> +               setup_git_directory(the_repository);
> +               repo_config(the_repository, git_default_config, NULL);
> +               prepare_repo_settings(the_repository);
> +               the_repository->settings.command_requires_full_index = 0;
> +
> +               repo_read_index(the_repository);
> +               ret = repo_index_has_changes(the_repository, NULL, &sb);
> +               printf("has_changes=%d\n", ret);
> +               if (sb.len)
> +                       printf("dirty=%s\n", sb.buf);

This seems to presume a single dirty file, otherwise wouldn't the
printing look pretty odd?

```

## Sahitya Chandra, 2026-08-07 08:05

Subject: Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD
Message-ID: <CAP=WS+sp74WQ=xndQ+2a6W-qP3Zz8=bVnEymgVpS+gwMv1Dh7g@mail.gmail.com>
In-Reply-To: <CABPp-BGYuQA_ngR3xS-_Mndzf_ubkn7rSc25CJG=UbLCVGdnyg@mail.gmail.com>

```
On Fri, Aug 7, 2026 at 12:17 PM Elijah Newren <newren@gmail.com> wrote:
>
> This explains what, but not why. It feels like a pedagogical exercise
> with no actual utility.

You are right, I found this through the TODO comment and do not have a
concrete user bug report or use case driving it.

> Why would someone with an unborn HEAD be using a sparse index? [...]

I do not have a good answer to that. My thinking was simply that
removing the special-case fallback still has some value: it deletes a
long-standing TODO, unifies the unborn-branch path with the normal diff
path, and removes an ensure_full_index() call that future readers would
need to reason about.

> This seems to presume a single dirty file, otherwise wouldn't the
> printing look pretty odd?

I agree that "dirty=%s" looks wrong when multiple paths are
present. I can fix that in v2.

Thanks for the review.

```

## Elijah Newren, 2026-08-07 15:33

Subject: Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD
Message-ID: <CABPp-BHLaW6_CxMdPQURN7zMK1p7dEkihFMAkyWvcd2+j7gJqw@mail.gmail.com>
In-Reply-To: <CAP=WS+sp74WQ=xndQ+2a6W-qP3Zz8=bVnEymgVpS+gwMv1Dh7g@mail.gmail.com>

```
On Fri, Aug 7, 2026 at 1:05 AM Sahitya Chandra <sahityajb@gmail.com> wrote:
>
> On Fri, Aug 7, 2026 at 12:17 PM Elijah Newren <newren@gmail.com> wrote:
> >
> > This explains what, but not why. It feels like a pedagogical exercise
> > with no actual utility.
>
> You are right, I found this through the TODO comment and do not have a
> concrete user bug report or use case driving it.
>
> > Why would someone with an unborn HEAD be using a sparse index? [...]
>
> I do not have a good answer to that. My thinking was simply that
> removing the special-case fallback still has some value: it deletes a
> long-standing TODO, unifies the unborn-branch path with the normal diff
> path, and removes an ensure_full_index() call that future readers would
> need to reason about.

Ah, thanks for looking through the code for TODOs and trying to clean
them up.  That's noble.

If you submit a v2, it's probably worth just being upfront about this
in the commit message -- that we don't expect this to be used in
practice, but it makes sense both (a) to remove one more TODO, and (b)
because it provides a net reduction in lines of code in read-cache.c.

> > This seems to presume a single dirty file, otherwise wouldn't the
> > printing look pretty odd?
>
> I agree that "dirty=%s" looks wrong when multiple paths are
> present. I can fix that in v2.

:-)

I'm curious if the unittesting harness could help here and avoid the
need for the test helper changes.  Is that possible?  (I don't
actually know much about the unittesting harness abilities, so I'm
genuinely curious).

> Thanks for the review.

Thanks for contributing!

```

## Sahitya Chandra, 2026-09-11 21:31

Subject: Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD
Message-ID: <CAP=WS+vySE94LZ_4CtEU6Hd99DttpF-49F3OaChTU53hdbzf1g@mail.gmail.com>
In-Reply-To: <CABPp-BHLaW6_CxMdPQURN7zMK1p7dEkihFMAkyWvcd2+j7gJqw@mail.gmail.com>

```
Hi Elijah,
Sorry for the long gap, and thanks for the review.

On Fri, Aug 7, 2026 at 9:03 PM Elijah Newren <newren@gmail.com> wrote:
>
> If you submit a v2, it's probably worth just being upfront about this
> in the commit message

That makes sense. I'll frame it as a cleanup of the TODO and a reduction
in read-cache.c, with no known practical use case.

> I'm curious if the unittesting harness could help here and avoid the
> need for the test helper changes. Is that possible?

I looked into the Clar harness and the existing unit tests. It supports
fixtures, but I didn't find repository/index setup helpers to reuse for
this case. A unit test may be possible with additional setup, though
keeping this in t1092 would let us reuse its sparse repository setup
and Trace2 checks to verify that the index stays unexpanded.

Would keeping the shell test and making the helper output clearer for
multiple paths be reasonable here?

Thanks,
Sahitya

```
