Volume XXII, number 279Tuesday, October 6, 2026Latest message 46 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

patchread-cache: avoid sparse-index expansion for unborn HEAD

6 messages between Aug 2, 2026 and Sep 11, 2026, from Sahitya Chandra, Elijah Newren.

Plain Markdown or JSON for tools and agents. Diffs are folded; open one to read it.

Sahitya ChandraAug 2, 2026, 21:28 UTC on lore

repo_index_has_changes() normally checks whether the index differs from a tree by passing that tree to the diff machinery. When no tree is passed, it tries to use HEAD for that comparison.

If HEAD does not resolve, as on an unborn branch, the function falls back to walking the index directly. With a sparse index, however, sparse directory entries may stand in for many paths, so the fallback first expands the index before reporting the changed paths.

That expansion is unnecessary. An unborn HEAD is equivalent for this check to comparing the index against the empty tree: every index entry is new relative to that tree.

Use the empty tree when HEAD cannot be resolved. This keeps the unborn-branch case on the same diff code path as the normal tree-comparison case, avoiding the sparse-index expansion while still letting callers see paths inside sparse directories.

Teach test-tool read-cache to exercise repo_index_has_changes(), and add a t1092 check that the unborn-branch case reports paths inside a sparse directory without expanding the index.

Signed-off-by: Sahitya Chandra <sahityajb@gmail.com>
---
 read-cache.c                             | 46 ++++++++++--------------
 t/helper/test-read-cache.c               | 19 ++++++++++
 t/t1092-sparse-checkout-compatibility.sh | 16 +++++++++
 3 files changed, 53 insertions(+), 28 deletions(-)
Show changes to 3 files +53 −28

read-cache.c, t/helper/test-read-cache.c, t/t1092-sparse-checkout-compatibility.sh

diff --git a/read-cache.c b/read-cache.c
index 6c449f393d..88ee9ba935 100644
--- a/read-cache.c
+++ b/read-cache.c
@@ -2505,39 +2505,29 @@ int repo_index_has_changes(struct repository *repo,
 			   struct tree *tree,
 			   struct strbuf *sb)
 {
-	struct index_state *istate = repo->index;
+	struct diff_options opt;
 	struct object_id cmp;
 	int i;
 
 	if (tree)
 		cmp = tree->object.oid;
-	if (tree || !repo_get_oid_tree(repo, "HEAD", &cmp)) {
-		struct diff_options opt;
-
-		repo_diff_setup(repo, &opt);
-		opt.flags.exit_with_status = 1;
-		if (!sb)
-			opt.flags.quick = 1;
-		diff_setup_done(&opt);
-		do_diff_cache(&cmp, &opt);
-		diffcore_std(&opt);
-		for (i = 0; sb && i < diff_queued_diff.nr; i++) {
-			if (i)
-				strbuf_addch(sb, ' ');
-			strbuf_addstr(sb, diff_queued_diff.queue[i]->two->path);
-		}
-		diff_flush(&opt);
-		return opt.flags.has_changes != 0;
-	} else {
-		/* TODO: audit for interaction with sparse-index. */
-		ensure_full_index(istate);
-		for (i = 0; sb && i < istate->cache_nr; i++) {
-			if (i)
-				strbuf_addch(sb, ' ');
-			strbuf_addstr(sb, istate->cache[i]->name);
-		}
-		return !!istate->cache_nr;
-	}
+	else if (repo_get_oid_tree(repo, "HEAD", &cmp))
+		oidcpy(&cmp, repo->hash_algo->empty_tree);
+
+	repo_diff_setup(repo, &opt);
+	opt.flags.exit_with_status = 1;
+	if (!sb)
+		opt.flags.quick = 1;
+	diff_setup_done(&opt);
+	do_diff_cache(&cmp, &opt);
+	diffcore_std(&opt);
+	for (i = 0; sb && i < diff_queued_diff.nr; i++) {
+		if (i)
+			strbuf_addch(sb, ' ');
+		strbuf_addstr(sb, diff_queued_diff.queue[i]->two->path);
+	}
+	diff_flush(&opt);
+	return opt.flags.has_changes != 0;
 }
 
 static int write_index_ext_header(struct hashfile *f,
diff --git a/t/helper/test-read-cache.c b/t/helper/test-read-cache.c
index 6b08ba8f07..ee629fbc69 100644
--- a/t/helper/test-read-cache.c
+++ b/t/helper/test-read-cache.c
@@ -4,6 +4,7 @@
 #include "config.h"
 #include "environment.h"
 #include "read-cache-ll.h"
+#include "repo-settings.h"
 #include "repository.h"
 #include "setup.h"
 
@@ -12,6 +13,24 @@ int cmd__read_cache(int argc, const char **argv)
 	int i, cnt = 1;
 	const char *name = NULL;
 
+	if (argc == 2 && !strcmp(argv[1], "--index-has-changes")) {
+		struct strbuf sb = STRBUF_INIT;
+		int ret;
+
+		setup_git_directory(the_repository);
+		repo_config(the_repository, git_default_config, NULL);
+		prepare_repo_settings(the_repository);
+		the_repository->settings.command_requires_full_index = 0;
+
+		repo_read_index(the_repository);
+		ret = repo_index_has_changes(the_repository, NULL, &sb);
+		printf("has_changes=%d\n", ret);
+		if (sb.len)
+			printf("dirty=%s\n", sb.buf);
+		strbuf_release(&sb);
+		return 0;
+	}
+
 	if (argc > 1 && skip_prefix(argv[1], "--print-and-refresh=", &name)) {
 		argc--;
 		argv++;
diff --git a/t/t1092-sparse-checkout-compatibility.sh b/t/t1092-sparse-checkout-compatibility.sh
index 4140c4d8ef..90239a862d 100755
--- a/t/t1092-sparse-checkout-compatibility.sh
+++ b/t/t1092-sparse-checkout-compatibility.sh
@@ -1558,6 +1558,22 @@ test_expect_success 'sparse-index is not expanded' '
 	)
 '
 
+test_expect_success 'sparse-index is not expanded: index has changes on unborn branch' '
+	init_repos &&
+	git -C sparse-index checkout --orphan unborn &&
+	git -C sparse-index ls-files --sparse --stage >cache &&
+	test_grep "^040000 .*	folder1/$" cache &&
+
+	rm -f trace2.txt &&
+	GIT_TRACE2_EVENT="$(pwd)/trace2.txt" GIT_TRACE2_EVENT_NESTING=10 \
+		test-tool -C sparse-index read-cache --index-has-changes \
+		>sparse-index-out 2>sparse-index-error &&
+	test_region ! index ensure_full_index trace2.txt &&
+	test_must_be_empty sparse-index-error &&
+	test_grep "has_changes=1" sparse-index-out &&
+	test_grep "folder1/a" sparse-index-out
+'
+
 test_expect_success 'sparse-index is not expanded: merge conflict in cone' '
 	init_repos &&
 

base-commit: a97fcc37c2bc6340a8d7ce78dedf227aac4e9aa7
-- 
2.43.0
Sahitya ChandraAug 6, 2026, 15:08 UTC in reply to Sahitya Chandra on lore

Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD

Just a gentle ping on this patch.

Thanks, Sahitya

Elijah NewrenAug 7, 2026, 06:47 UTC in reply to Sahitya Chandra on lore

Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD

On Sun, Aug 2, 2026 at 2:28 PM Sahitya Chandra <sahityajb@gmail.com> wrote:
Show 22 quoted lines
>
> repo_index_has_changes() normally checks whether the index differs from
> a tree by passing that tree to the diff machinery. When no tree is
> passed, it tries to use HEAD for that comparison.
>
> If HEAD does not resolve, as on an unborn branch, the function falls
> back to walking the index directly. With a sparse index, however, sparse
> directory entries may stand in for many paths, so the fallback first
> expands the index before reporting the changed paths.
>
> That expansion is unnecessary. An unborn HEAD is equivalent for this
> check to comparing the index against the empty tree: every index entry
> is new relative to that tree.
>
> Use the empty tree when HEAD cannot be resolved. This keeps the
> unborn-branch case on the same diff code path as the normal
> tree-comparison case, avoiding the sparse-index expansion while still
> letting callers see paths inside sparse directories.
>
> Teach test-tool read-cache to exercise repo_index_has_changes(), and
> add a t1092 check that the unborn-branch case reports paths inside a
> sparse directory without expanding the index.

This explains what, but not why. It feels like a pedagogical exercise with no actual utility. Why would someone with an unborn HEAD be using a sparse index? They have millions of files, with none of them committed, except they don't have millions of files because they only have paths under certain directories? How did they even get the relevant tree entries into the sparse index in order to have one?

Perhaps you have a great usecase and I've just missed it. Could you explain the motivation for enabling this? Or was it more a case of trying to take care of TODOs in the code?

[...]
Show 18 quoted lines
> @@ -12,6 +13,24 @@ int cmd__read_cache(int argc, const char **argv)
>         int i, cnt = 1;
>         const char *name = NULL;
>
> +       if (argc == 2 && !strcmp(argv[1], "--index-has-changes")) {
> +               struct strbuf sb = STRBUF_INIT;
> +               int ret;
> +
> +               setup_git_directory(the_repository);
> +               repo_config(the_repository, git_default_config, NULL);
> +               prepare_repo_settings(the_repository);
> +               the_repository->settings.command_requires_full_index = 0;
> +
> +               repo_read_index(the_repository);
> +               ret = repo_index_has_changes(the_repository, NULL, &sb);
> +               printf("has_changes=%d\n", ret);
> +               if (sb.len)
> +                       printf("dirty=%s\n", sb.buf);

This seems to presume a single dirty file, otherwise wouldn't the printing look pretty odd?

Sahitya ChandraAug 7, 2026, 08:05 UTC in reply to Elijah Newren on lore

Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD

On Fri, Aug 7, 2026 at 12:17 PM Elijah Newren <newren@gmail.com> wrote:
>
> This explains what, but not why. It feels like a pedagogical exercise
> with no actual utility.

You are right, I found this through the TODO comment and do not have a concrete user bug report or use case driving it.

> Why would someone with an unborn HEAD be using a sparse index? [...]

I do not have a good answer to that. My thinking was simply that removing the special-case fallback still has some value: it deletes a long-standing TODO, unifies the unborn-branch path with the normal diff path, and removes an ensure_full_index() call that future readers would need to reason about.

> This seems to presume a single dirty file, otherwise wouldn't the
> printing look pretty odd?

I agree that "dirty=%s" looks wrong when multiple paths are present. I can fix that in v2.

Thanks for the review.
Elijah NewrenAug 7, 2026, 15:33 UTC in reply to Sahitya Chandra on lore

Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD

On Fri, Aug 7, 2026 at 1:05 AM Sahitya Chandra <sahityajb@gmail.com> wrote:
Show 16 quoted lines
>
> On Fri, Aug 7, 2026 at 12:17 PM Elijah Newren <newren@gmail.com> wrote:
> >
> > This explains what, but not why. It feels like a pedagogical exercise
> > with no actual utility.
>
> You are right, I found this through the TODO comment and do not have a
> concrete user bug report or use case driving it.
>
> > Why would someone with an unborn HEAD be using a sparse index? [...]
>
> I do not have a good answer to that. My thinking was simply that
> removing the special-case fallback still has some value: it deletes a
> long-standing TODO, unifies the unborn-branch path with the normal diff
> path, and removes an ensure_full_index() call that future readers would
> need to reason about.

Ah, thanks for looking through the code for TODOs and trying to clean them up. That's noble.

If you submit a v2, it's probably worth just being upfront about this in the commit message -- that we don't expect this to be used in practice, but it makes sense both (a) to remove one more TODO, and (b) because it provides a net reduction in lines of code in read-cache.c.

Show 5 quoted lines
> > This seems to presume a single dirty file, otherwise wouldn't the
> > printing look pretty odd?
>
> I agree that "dirty=%s" looks wrong when multiple paths are
> present. I can fix that in v2.
:-)

I'm curious if the unittesting harness could help here and avoid the need for the test helper changes. Is that possible? (I don't actually know much about the unittesting harness abilities, so I'm genuinely curious).

> Thanks for the review.
Thanks for contributing!
Sahitya ChandraSep 11, 2026, 21:31 UTC in reply to Elijah Newren on lore

Re: [PATCH] read-cache: avoid sparse-index expansion for unborn HEAD

Hi Elijah, Sorry for the long gap, and thanks for the review.

On Fri, Aug 7, 2026 at 9:03 PM Elijah Newren <newren@gmail.com> wrote:
>
> If you submit a v2, it's probably worth just being upfront about this
> in the commit message

That makes sense. I'll frame it as a cleanup of the TODO and a reduction in read-cache.c, with no known practical use case.

> I'm curious if the unittesting harness could help here and avoid the
> need for the test helper changes. Is that possible?

I looked into the Clar harness and the existing unit tests. It supports fixtures, but I didn't find repository/index setup helpers to reuse for this case. A unit test may be possible with additional setup, though keeping this in t1092 would let us reuse its sparse repository setup and Trace2 checks to verify that the index stays unexpanded.

Would keeping the shell test and making the helper output clearer for multiple paths be reasonable here?

Thanks, Sahitya

Back to recent threads