[PATCH v2 0/3] ls-files: reuse and update the untracked cache
- From
Tamir Duberstein <tamird@gmail.com>
- Date
- Sep 23, 2026, 22:26 UTC
- Message-ID
- <20260923-ls-files-untracked-cache-v2-0-d7ee33476eb8@gmail.com>
- In-Reply-To
- <20260923-ls-files-untracked-cache-v1-0-08db4cc1efdb@gmail.com>
Repeated queries such as
git ls-files --cached --others --exclude-standard -z -- "**/pyproject.toml"
walk the working tree even when status has populated an untracked cache. This series lets ls-files reuse those listings and save the work for later commands through Git's existing optional index writes.
The first patch fixes inconsistent ignore-file hashes that invalidate an unchanged cache. The second lets 'git status -unormal' and 'git status -uall' reuse the same cache. 'git status -unormal' still stops scanning an untracked directory after finding an untracked file; a later command requesting all untracked files completes the listing as needed. The third lets ls-files use the cache and write pending untracked cache and fsmonitor updates back to the index. --no-optional-locks suppresses those writes.
For wildcard pathspecs without a fixed prefix, apply the pathspec after reading complete directory listings. A first query can therefore scan directories that its pathspec would otherwise skip. Fixed-prefix, attribute and exclude pathspecs retain their existing traversal.
A new dir_flags value makes older versions of Git rebuild the cache before using it to list untracked files. Changes invalidate cached entries in parent directories as well, so updates can reopen more directories than a cache populated with --untracked-files=all before this series.
On macOS, with 100,000 files in 5,000 leaf directories, half tracked. Each group of 100 leaf directories has 50 tracked and 50 untracked. Each binary populated its own cache using the indicated status mode.
hyperfine --warmup 3 --runs 7 (mean ± standard deviation, milliseconds):
cache fsmonitor command upstream v1 v2 normal off query 241.3 ± 14.8 132.7 ± 20.0 33.3 ± 2.4 normal off query + status 292.7 ± 11.2 187.0 ± 10.8 92.6 ± 3.3 all off query 231.3 ± 14.6 29.3 ± 1.3 34.7 ± 2.0 normal on query 223.1 ± 17.8 111.0 ± 13.4 33.8 ± 1.5 all on query 216.6 ± 12.0 28.0 ± 5.1 34.6 ± 0.8
"query" is the ls-files command above; "status" is "git status -unormal --porcelain". These are repeated queries with warm filesystem and untracked caches.
Assisted-by: LLM Signed-off-by: Tamir Duberstein <tamird@gmail.com> --- Changes in v2: - Share the cache between --untracked-files=normal and --untracked-files=all. - Write ls-files untracked cache and fsmonitor updates back to the index. - Test alternating modes, cache invalidation and optional index writes. - Replace the v1 measurements with results for the revised implementation. - Link to v1: https://patch.msgid.link/20260923-ls-files-untracked-cache-v1-0-08db4cc1efdb@gmail.com
---
Tamir Duberstein (3):
dir: hash ignore files before appending newline
dir: share untracked caches across output modes
ls-files: use and update the untracked cacheDocumentation/git-ls-files.adoc | 4 + Documentation/gitformat-index.adoc | 13 +- builtin/ls-files.c | 40 +++++- dir.c | 251 ++++++++++++++++++---------------- dir.h | 16 +-- t/perf/p3010-ls-files.sh | 15 ++ t/t3001-ls-files-others-exclude.sh | 20 +++ t/t7063-status-untracked-cache.sh | 272 +++++++++++++++++++++++++++++-------- t/t7519-status-fsmonitor.sh | 44 ++++++ 9 files changed, 481 insertions(+), 194 deletions(-)
Range-diff versus v1:
1: 896f1aea52 ! 1: bf5e5eea24 dir: hash ignore files before adding parser LF
@@ Metadata
Author: Tamir Duberstein <tamird@gmail.com>
## Commit message ##
- dir: hash ignore files before adding parser LF
+ dir: hash ignore files before appending newline
add_patterns() appends a newline for the pattern parser before computing
- an ignore file's object ID. Its fallback hash therefore includes a byte
- that is absent from the file. The fast path instead copies the original
- blob ID from an up-to-date index entry.
+ an ignore file's object ID. Hashing the buffer therefore includes a byte
+ that is absent from the file. When the file has an up-to-date index entry
+ and needs no content conversion, the function instead uses that entry's
+ object ID.
- Switching between those paths changes the recorded ignore identity even
- when the file has not changed, invalidating the untracked cache below it.
- Compute the hash before appending the parser newline so both paths agree.
- Update the expected identities of the untracked ignore files accordingly.
+ Switching between these paths changes the cached object ID even when the
+ file has not changed, invalidating the untracked cache below it. Compute
+ the hash before appending the newline so both paths agree. Update the
+ expected object IDs of the untracked ignore files accordingly.
+ Assisted-by: LLM
Signed-off-by: Tamir Duberstein <tamird@gmail.com>
## dir.c ##
@@ dir.c: static int add_patterns(const char *fname, const char *base, int baselen,
fill_stat_data(&oid_stat->stat, &st);
oid_stat->valid = 1;
}
++ /*
++ * The extra newline is only for parsing. Like do_read_blob(),
++ * keep it out of the file's object ID.
++ */
+ buf[size++] = '\n';
}
2: fc65e91309 < -: ---------- ls-files: reuse cached untracked listings
-: ---------- > 2: ae86ca623b dir: share untracked caches across output modes
-: ---------- > 3: 126241f862 ls-files: use and update the untracked cache--- base-commit: 3bc0341126508f78f5869cbfc0005e987efdf0c7 change-id: 20260923-ls-files-untracked-cache-3559bed01a3e