{"thread":{"id":"66492","subject":"[PATCH 0/2] status: agree with diff and add when conversion is active","startedAt":"2026-10-08T20:45:02Z","lastAt":"2026-10-09T05:37:27Z","messageCount":4,"participants":["Curtis Allen Smith","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":2},"messages":[{"id":"554534","messageId":"20261008204603.1988-1-curtis.allen.smith@gmail.com","threadId":"66492","inReplyTo":null,"subject":"[PATCH 0/2] status: agree with diff and add when conversion is active","fromName":"Curtis Allen Smith","fromEmail":"curtis.allen.smith@gmail.com","sentAt":"2026-10-08T20:45:02Z","receivedAt":"2026-10-08T20:45:02Z","isPatch":true,"sender":{"key":"curtis.allen.smith@gmail.com","avatar":null},"body":"\"git status\", \"git diff\" and \"git add\" can disagree about whether a\nfile has been modified.  Under\n\n\t* text eol=lf\n\na tool that rewrites an otherwise unchanged file with CRLF endings\nmakes \"git status\" report it as modified, while \"git diff\" shows\nnothing and \"git add\" stages nothing: the two commands that actually\nrun the clean filter both conclude that the contents did not change.\nThat contradiction, rather than the line endings as such, is what\nthis series is about.\n\nThe cause is the size comparison in ie_match_stat().  ie_modified()\ntakes a difference between the file's size and the size recorded in\nthe index as proof of a content change and returns without reading\nthe file.  That was sound in 2005, when the working tree file and the\nblob were the same bytes.  Conversion made it unsound -- the whole\npoint of a clean filter is that the two representations differ in\ntheir bytes and agree on their content -- and the shortcut was never\nrevisited.  The mtime branch of the very same function already reads\nthe file and applies the conversion before deciding, so Git pays for\nthe conversion-aware check in one branch and refuses to in the other.\n\nPatch 1 makes the size branch behave like the mtime branch whenever\nthe path is subject to conversion.  Paths with no conversion take the\nexisting early return untouched.  It also stops ce_compare_data()\nhashing a file whose converted length already differs from the size\nof its blob, since equal contents must have equal length; that is\nwhere most of the cost of the new check would otherwise go, and it\nhelps the pre-existing mtime path as well.  When the end-of-line\nconversion is the only one that applies, that length comes from a\nscan for CR, and the file is not converted either.\n\nPatch 2 adds core.convertAwareStatus for people who would rather keep\nthe old shortcut, either everywhere (false) or only for paths with an\nexpensive clean filter such as Git LFS (no-filter).  I defaulted it to\non, including filters: the measurements below say reading is not what\ncosts, and the 2005 performance argument should not be re-applied in\n2026 without evidence.  Being ordinary configuration it also works per\ncommand, as \"git -c core.convertAwareStatus=no-filter status\".\n\nNote that there is no way to opt out of the size comparison today --\ncore.checkStat=minimal drops ctime, uid/gid and inode but still\ncompares the size -- which is why patch 2 adds a variable instead of\nextending an existing one.\n\nNumbers\n-------\n\nLinux (WSL2, ext4) on an i9-14900K, warm page cache, fastest of 5\nruns, \"status -uno\" to separate the refresh from untracked scanning.\nOne binary for both columns with core.convertAwareStatus flipped;\n\"false\" is the pre-series code path.\n\n\t10000 files x 2.6 KB, \"* text=auto eol=lf\"\n\t                                   false      true\n\t  clean tree                         4 ms      4 ms\n\t  10000 genuinely modified          19 ms     49 ms\n\t  10000 CRLF-rewritten, 1st run     20 ms    148 ms\n\t  the same, steady state            20 ms      6 ms\n\n\t200 files x 1 MB, same attributes\n\t                                   false      true\n\t  clean tree                         2 ms      2 ms\n\t  200 genuinely modified             2 ms     26 ms\n\t  200 CRLF-rewritten, 1st run        2 ms    611 ms\n\t  the same, steady state             2 ms      3 ms\n\n\t10000 files x 2.6 KB, no conversion configured\n\t  clean tree                         4 ms      4 ms\n\t  10000 genuinely modified          19 ms     20 ms\n\nA clean tree and a repository without conversion are unaffected.  What\nis paid for is stat-dirty converted paths.  A file that was really\nedited is read and scanned for CR, but neither converted nor hashed,\nbecause its length already rules out a match.  Without that the\n\"genuinely modified\" rows read 142 ms and 517 ms rather than 49 ms and\n26 ms.  Of the 214 MB in the second corpus, reading costs 6 ms from\npage cache, the conversion about 180 ms, and SHA-1 about 280 ms.\n\nThe last row of the first block is the case the series exists for: the\npatched build settles at 6 ms where the unpatched one pays 20 ms on\nevery invocation and still reports the files as modified, because it\nnever refreshes their recorded sizes.  The 1st-run rows are the\none-time cost of discovering that.  Their lengths match, so they are\nconverted and hashed in full.\n\nThis was reported against Git for Windows [1], where Torsten suggested\nbringing it to the list.  It is not Windows-specific; anything with a\nclean filter runs into it, and Git LFS users on Linux see the same\ncontradiction.\n\nBuilt with gcc 15.2 on top of 6de20f6; each commit builds and passes\non its own.  t0020 (with the new tests), t0021, t0026, t0027, t1300,\nt2106, t2200, t3700, t7508 and t0008 pass.\n\n[1] https://github.com/git-for-windows/git/issues/6410\n\nCurtis Allen Smith (2):\n  read-cache: do not trust a size change when conversion is active\n  core: add core.convertAwareStatus to opt out of the content check\n\n Documentation/config/core.adoc |  22 +++++\n environment.c                  |  14 ++++\n environment.h                  |   7 ++\n read-cache.c                   | 141 ++++++++++++++++++++++++++++++++-\n t/t0020-crlf.sh                |  92 +++++++++++++++++++++\n 5 files changed, 273 insertions(+), 3 deletions(-)\n\n-- \n2.53.0\n\n\n"},{"id":"554535","messageId":"20261008204603.1988-2-curtis.allen.smith@gmail.com","threadId":"66492","inReplyTo":"20261008204603.1988-1-curtis.allen.smith@gmail.com","subject":"[PATCH 1/2] read-cache: do not trust a size change when conversion is active","fromName":"Curtis Allen Smith","fromEmail":"curtis.allen.smith@gmail.com","sentAt":"2026-10-08T20:45:03Z","receivedAt":"2026-10-08T20:45:03Z","isPatch":true,"sender":{"key":"curtis.allen.smith@gmail.com","avatar":null},"body":"\"git status\" can report a file as modified while \"git diff\" and\n\"git add\", which both run the clean filter, agree its contents are\nunchanged:\n\n\tgit init t && cd t\n\tprintf '* text eol=lf\\n' >.gitattributes\n\tprintf 'one\\ntwo\\nthree\\n' >file.txt\n\tgit add . && git commit -m init\n\n\tprintf 'one\\r\\ntwo\\r\\nthree\\r\\n' >file.txt\n\n\tgit status --short      # ->  M file.txt\n\tgit diff                # -> empty\n\tgit add file.txt        # -> stages nothing\n\nThree commands that answer the same question answer it differently,\nand the two that consult the conversion are the ones that get it\nright.\n\nie_modified() declares a path modified as soon as ie_match_stat()\nreports DATA_CHANGED, which it does whenever the size of the file in\nthe working tree differs from the size recorded in the index, and it\nreturns without ever reading the file.  That shortcut is sound only\nwhile the working tree file and the blob are the same bytes.  When it\nwas written in 2005 they were, and a size mismatch really was a proof\nof a content change.  Conversion removed that premise: core.autocrlf\narrived in 2007 and the \"text\" and \"eol\" attributes in 2010, and\nmaking the two representations differ in their bytes while agreeing on\ntheir content is precisely what they are for.  The shortcut was never\nre-examined against the feature layered on top of it.\n\nThe same function already handles the comparable case correctly.  When\nonly the mtime changed, it falls through to ce_modified_check_fs(),\nreads the file, applies the conversion, and answers \"unchanged\" when\nthat is the truth.  Git is therefore already willing to pay for a\nconversion-aware check here; the size branch is the only place where a\ndifference in the bytes on disk is taken to be a difference in\ncontent.\n\nSo fall through to that same check when the size changed and the path\nis subject to conversion.  Paths without conversion take the early\nreturn exactly as before, and a repository that uses no conversion is\nunaffected.\n\nHashing is the expensive part of that check -- Git's collision\ndetecting SHA-1 runs at about 800 MB/s on the machine used below --\nand it is avoidable most of the time.  Contents that are equal\nnecessarily have equal length, so ce_compare_data() now compares the\nlength of the converted file against the size of the blob, which costs\nan object header lookup, and hashes only when the two agree.  A file\nthat was really edited almost always changes length and is rejected\nwithout being hashed.  The file this commit is about has exactly the\nlength of its blob, so it is hashed, found equal, and the index then\nrecords its new size, after which it is not read again.\n\nWhen the end-of-line conversion is the only one that applies, even\nthe conversion can be skipped.  \"git add\" either keeps such a file as\nit is or turns each CRLF into LF, so the converted length is one of\ntwo numbers, and a scan for CR gives both.  If neither is the size of\nthe blob, the file is modified, and it is neither converted nor\nhashed.\n\nIn a repository of 200 files of 1 MB each under \"* text=auto\" with\nevery file modified, these take \"git status\" from 517ms to 26ms; with\n10000 files of 2.6 KB, from 142ms to 49ms.\n\nThe inconsistency is a chronic annoyance for anyone sharing a tree\nbetween Windows and Unix with normalized line endings.  Any tool that\nrewrites unchanged files in native line endings -- javadoc, code\ngenerators, formatters, a good number of editors -- changes the size\nof every file it touches and flags the whole output tree as modified\nwith empty diffs.  The documented remedy, \"git add --renormalize\",\ndoes not stick: the next checkout, stash or branch switch rewrites the\ncheckout-form bytes and the recorded sizes along with them, and the\nnext run of the tool flags everything again.\n\nSigned-off-by: Curtis Allen Smith <curtis.allen.smith@gmail.com>\n---\n read-cache.c    | 129 ++++++++++++++++++++++++++++++++++++++++++++++--\n t/t0020-crlf.sh |  45 +++++++++++++++++\n 2 files changed, 171 insertions(+), 3 deletions(-)\n\ndiff --git a/read-cache.c b/read-cache.c\nindex c4cf08a3a..8875706d8 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -9,6 +9,7 @@\n \n #include \"git-compat-util.h\"\n #include \"config.h\"\n+#include \"convert.h\"\n #include \"date.h\"\n #include \"diff.h\"\n #include \"diffcore.h\"\n@@ -228,6 +229,99 @@ int fake_lstat(const struct cache_entry *ce, struct stat *st)\n \treturn 0;\n }\n \n+/*\n+ * Count the CRs in \"buf\" that are immediately followed by an LF.\n+ */\n+static size_t count_crlf(const char *buf, size_t len)\n+{\n+\tconst char *end = buf + len;\n+\tsize_t n = 0;\n+\n+\twhile ((buf = memchr(buf, '\\r', end - buf))) {\n+\t\tif (++buf < end && *buf == '\\n')\n+\t\t\tn++;\n+\t}\n+\treturn n;\n+}\n+\n+/*\n+ * Compare an open file to the blob recorded for it, converting the file\n+ * the way \"git add\" would.  Contents that are equal necessarily have\n+ * equal length, so when the converted length differs from the size of\n+ * the blob the file is modified and there is no need to hash it, and\n+ * hashing is by far the most expensive part of this comparison.\n+ *\n+ * Only the case this can help is handled here: a regular file whose\n+ * conversion Git performs itself.  A path driven by an external filter\n+ * is left to index_fd(), which streams it into the filter.\n+ *\n+ * Returns 1 if the file differs, 0 if it matches, -1 if it could not be\n+ * read.  Does not close \"fd\".\n+ */\n+static int ce_compare_converted_data(struct index_state *istate,\n+\t\t\t\t     const struct cache_entry *ce,\n+\t\t\t\t     struct stat *st, int fd)\n+{\n+\tstruct strbuf raw = STRBUF_INIT;\n+\tstruct strbuf converted = STRBUF_INIT;\n+\tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct object_id oid;\n+\tstruct conv_attrs ca;\n+\tenum object_type type;\n+\tsize_t blob_size;\n+\tconst char *buf;\n+\tsize_t len;\n+\tint have_size, match = -1;\n+\n+\toi.typep = &type;\n+\toi.sizep = &blob_size;\n+\thave_size = (odb_read_object_info_extended(istate->repo->objects,\n+\t\t\t\t\t\t   &ce->oid, &oi,\n+\t\t\t\t\t\t   OBJECT_INFO_SKIP_FETCH_OBJECT |\n+\t\t\t\t\t\t   OBJECT_INFO_QUICK) == ODB_READ_OK &&\n+\t\t     type == OBJ_BLOB);\n+\n+\tif (strbuf_read(&raw, fd, st->st_size) < 0)\n+\t\tgoto out;\n+\n+\t/*\n+\t * When the end-of-line conversion is the only one, \"git add\"\n+\t * either keeps the file as it is or turns every CRLF into LF\n+\t * (\"text=auto\" refuses to convert a file with a lone CR, so\n+\t * stripping all CRs comes to the same thing).  The converted\n+\t * length is therefore one of two values, and if neither is the\n+\t * size of the blob the file is modified without converting it.\n+\t */\n+\tconvert_attrs(istate, &ca, ce->name);\n+\tif (have_size && !ca.drv && !ca.ident &&\n+\t    !ca.working_tree_encoding &&\n+\t    raw.len != blob_size &&\n+\t    raw.len - count_crlf(raw.buf, raw.len) != blob_size) {\n+\t\tmatch = 1;\n+\t\tgoto out;\n+\t}\n+\n+\tbuf = raw.buf;\n+\tlen = raw.len;\n+\tif (convert_to_git(istate, ce->name, raw.buf, raw.len, &converted, 0)) {\n+\t\tbuf = converted.buf;\n+\t\tlen = converted.len;\n+\t}\n+\n+\tif (have_size && len != blob_size) {\n+\t\tmatch = 1;\n+\t\tgoto out;\n+\t}\n+\n+\thash_object_file(istate->repo->hash_algo, buf, len, OBJ_BLOB, &oid);\n+\tmatch = !oideq(&oid, &ce->oid);\n+\n+out:\n+\tstrbuf_release(&raw);\n+\tstrbuf_release(&converted);\n+\treturn match;\n+}\n+\n static int ce_compare_data(struct index_state *istate,\n \t\t\t   const struct cache_entry *ce,\n \t\t\t   struct stat *st)\n@@ -237,9 +331,16 @@ static int ce_compare_data(struct index_state *istate,\n \n \tif (fd >= 0) {\n \t\tstruct object_id oid;\n-\t\tif (!index_fd(istate, &oid, fd, st, OBJ_BLOB, ce->name, 0))\n+\n+\t\tif (S_ISREG(st->st_mode) &&\n+\t\t    would_convert_to_git(istate, ce->name) &&\n+\t\t    !would_convert_to_git_filter_fd(istate, ce->name)) {\n+\t\t\tmatch = ce_compare_converted_data(istate, ce, st, fd);\n+\t\t\tclose(fd);\n+\t\t} else if (!index_fd(istate, &oid, fd, st, OBJ_BLOB, ce->name, 0)) {\n \t\t\tmatch = !oideq(&oid, &ce->oid);\n-\t\t/* index_fd() closed the file descriptor already */\n+\t\t\t/* index_fd() closed the file descriptor already */\n+\t\t}\n \t}\n \treturn match;\n }\n@@ -438,6 +539,27 @@ int ie_match_stat(struct index_state *istate,\n \treturn changed;\n }\n \n+/*\n+ * A difference between the size of the file in the working tree and the\n+ * size recorded for it in the index proves that the contents changed\n+ * only as long as the two are byte-for-byte comparable.  That stops\n+ * being true as soon as the path is run through a clean filter:\n+ * rewriting a file with CRLF endings under \"text eol=lf\", for example,\n+ * changes its size in the working tree without changing the blob Git\n+ * would record for it.  For such a path the only way to tell is to read\n+ * the contents and convert them, which is what we already do when only\n+ * the mtime changed.\n+ */\n+static int size_change_is_conclusive(struct index_state *istate,\n+\t\t\t\t     const struct cache_entry *ce,\n+\t\t\t\t     struct stat *st)\n+{\n+\tif (!S_ISREG(st->st_mode))\n+\t\treturn 1;\n+\n+\treturn !would_convert_to_git(istate, ce->name);\n+}\n+\n int ie_modified(struct index_state *istate,\n \t\tconst struct cache_entry *ce,\n \t\tstruct stat *st, unsigned int options)\n@@ -480,7 +602,8 @@ int ie_modified(struct index_state *istate,\n \t     */\n \t    (!S_ISLNK(st->st_mode) || ce->ce_stat_data.sd_size != MAX_PATH) &&\n #endif\n-\t    (S_ISGITLINK(ce->ce_mode) || ce->ce_stat_data.sd_size != 0))\n+\t    (S_ISGITLINK(ce->ce_mode) || ce->ce_stat_data.sd_size != 0) &&\n+\t    size_change_is_conclusive(istate, ce, st))\n \t\treturn changed;\n \n \tchanged_fs = ce_modified_check_fs(istate, ce, st);\ndiff --git a/t/t0020-crlf.sh b/t/t0020-crlf.sh\nindex fd1cae09e..88b728d50 100755\n--- a/t/t0020-crlf.sh\n+++ b/t/t0020-crlf.sh\n@@ -397,4 +397,49 @@ test_expect_success 'New CRLF file gets LF in repo' '\n \ttest_cmp alllf alllf2\n '\n \n+test_expect_success 'status does not report a CRLF-only rewrite as modified' '\n+\tgit init eol-status &&\n+\t(\n+\t\tcd eol-status &&\n+\t\techo \"* text eol=lf\" >.gitattributes &&\n+\t\tprintf \"one\\ntwo\\nthree\\n\" >file.txt &&\n+\t\tgit add .gitattributes file.txt &&\n+\t\tgit commit -m initial &&\n+\n+\t\t# a generator rewrites the file with CRLF, same content\n+\t\tprintf \"one\\r\\ntwo\\r\\nthree\\r\\n\" >file.txt &&\n+\t\tgit status --porcelain -uno >actual &&\n+\t\ttest_must_be_empty actual &&\n+\t\tgit diff --exit-code &&\n+\n+\t\t# a real change is still reported\n+\t\tprintf \"one\\r\\ntwo\\r\\nfour\\r\\n\" >file.txt &&\n+\t\tgit status --porcelain -uno >actual &&\n+\t\techo \" M file.txt\" >expect &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'status sizes a text file by its CRLF pairs, not its CRs' '\n+\tgit init eol-status-lone-cr &&\n+\t(\n+\t\tcd eol-status-lone-cr &&\n+\t\techo \"* text eol=lf\" >.gitattributes &&\n+\t\tprintf \"one\\rtwo\\nthree\\n\" >file.txt &&\n+\t\tgit add .gitattributes file.txt &&\n+\t\tgit commit -m initial &&\n+\n+\t\t# \"git add\" keeps the lone CR and drops the others\n+\t\tprintf \"one\\rtwo\\r\\nthree\\r\\n\" >file.txt &&\n+\t\tgit status --porcelain -uno >actual &&\n+\t\ttest_must_be_empty actual &&\n+\n+\t\t# the converted length matches the blob, the content does not\n+\t\tprintf \"one\\rtwo\\nthrEE\\r\\n\" >file.txt &&\n+\t\tgit status --porcelain -uno >actual &&\n+\t\techo \" M file.txt\" >expect &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \n2.53.0\n\n\n"},{"id":"554536","messageId":"20261008204603.1988-3-curtis.allen.smith@gmail.com","threadId":"66492","inReplyTo":"20261008204603.1988-1-curtis.allen.smith@gmail.com","subject":"[PATCH 2/2] core: add core.convertAwareStatus to opt out of the content check","fromName":"Curtis Allen Smith","fromEmail":"curtis.allen.smith@gmail.com","sentAt":"2026-10-08T20:45:04Z","receivedAt":"2026-10-08T20:45:04Z","isPatch":true,"sender":{"key":"curtis.allen.smith@gmail.com","avatar":null},"body":"The previous commit makes an index refresh read and convert a path\nwhose size changed when conversion is active for it, so that \"git\nstatus\" agrees with \"git diff\" and \"git add\".  Reading costs more than\ntrusting the size, and when the path has a clean filter configured the\ncost includes running that filter -- Git LFS on a large file, say.\n\nReading is cheap enough on current hardware that agreeing with\n\"git diff\" is the better default, but nobody should be stuck with it\nif their filters are expensive.  Add core.convertAwareStatus:\n\n\ttrue (default)  consult the conversion for any path that has one,\n\t                including paths with a clean filter\n\tno-filter       consult only the conversions Git performs itself,\n\t                and decide a path with a clean filter on its size\n\tfalse           always treat a size change as a modification, as\n\t                Git did before\n\nBeing ordinary configuration, it can equally be given for a single\ncommand:\n\n\tgit -c core.convertAwareStatus=no-filter status\n\nPaths that are not subject to conversion are decided on their size\nalone in every mode, so this costs nothing in a repository that does\nnot use conversion.\n\nSigned-off-by: Curtis Allen Smith <curtis.allen.smith@gmail.com>\n---\n Documentation/config/core.adoc | 22 ++++++++++++++++\n environment.c                  | 14 ++++++++++\n environment.h                  |  7 +++++\n read-cache.c                   | 12 +++++++++\n t/t0020-crlf.sh                | 47 ++++++++++++++++++++++++++++++++++\n 5 files changed, 102 insertions(+)\n\ndiff --git a/Documentation/config/core.adoc b/Documentation/config/core.adoc\nindex 0b697f53f..9737c804f 100644\n--- a/Documentation/config/core.adoc\n+++ b/Documentation/config/core.adoc\n@@ -156,6 +156,28 @@ some fields (e.g. JGit); by excluding these fields from the\n comparison, the `minimal` mode may help interoperability when the\n same repository is used by these other systems at the same time.\n \n+core.convertAwareStatus::\n+\tWhen a path is subject to content conversion -- the `text` and\n+\t`eol` attributes, `core.autocrlf`, a `working-tree-encoding`,\n+\tor a clean filter -- the size of the file in the working tree\n+\tis not determined by its contents alone, so a change in size\n+\tdoes not prove that the contents changed.  When this variable\n+\tis missing or set to `true`, Git reads and converts such a\n+\tfile before reporting it as modified, which keeps 'git status'\n+\tin agreement with 'git diff' and 'git add'.  When set to\n+\t`no-filter`, Git does this only for the conversions it\n+\tperforms itself, and a path with a clean filter configured\n+\t(Git LFS, for example) is reported as modified on a size\n+\tchange without running the filter.  When set to `false`, a\n+\tsize change is always taken as a modification, which is what\n+\tGit did before this variable existed.\n++\n+Reading the file costs more than trusting its size, so `no-filter`\n+and `false` trade this consistency for speed in repositories where\n+running the filter, or reading the file at all, is too expensive.\n+Paths that are not subject to conversion are decided on the size\n+alone in every mode.\n+\n core.quotePath::\n \tCommands that output paths (e.g. 'ls-files', 'diff'), will\n \tquote \"unusual\" characters in the pathname by enclosing the\ndiff --git a/environment.c b/environment.c\nindex c83cf4483..b079262c8 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -344,6 +344,19 @@ int git_default_core_config(const char *var, const char *value,\n \t\t\t\t     var, value);\n \t}\n \n+\tif (!strcmp(var, \"core.convertawarestatus\")) {\n+\t\tint b = git_parse_maybe_bool(value);\n+\t\tif (0 <= b)\n+\t\t\tcfg->convert_aware_status = b ? CONVERT_AWARE_STATUS_ALL\n+\t\t\t\t\t\t      : CONVERT_AWARE_STATUS_NEVER;\n+\t\telse if (value && !strcasecmp(value, \"no-filter\"))\n+\t\t\tcfg->convert_aware_status = CONVERT_AWARE_STATUS_IN_PROCESS;\n+\t\telse\n+\t\t\treturn error(_(\"invalid value for '%s': '%s'\"),\n+\t\t\t\t     var, value);\n+\t\treturn 0;\n+\t}\n+\n \tif (!strcmp(var, \"core.quotepath\")) {\n \t\tquote_path_fully = git_config_bool(var, value);\n \t\treturn 0;\n@@ -766,6 +779,7 @@ void repo_config_values_init(struct repo_config_values *cfg)\n \tcfg->apply_sparse_checkout = 0;\n \tcfg->trust_ctime = 1;\n \tcfg->check_stat = 1;\n+\tcfg->convert_aware_status = CONVERT_AWARE_STATUS_ALL;\n \tcfg->zlib_compression_level = Z_BEST_SPEED;\n \tcfg->pack_compression_level = Z_DEFAULT_COMPRESSION;\n \tcfg->precomposed_unicode = -1; /* see probe_utf8_pathname_composition() */\ndiff --git a/environment.h b/environment.h\nindex b336459e9..49a88de27 100644\n--- a/environment.h\n+++ b/environment.h\n@@ -115,6 +115,12 @@ enum object_creation_mode {\n \tOBJECT_CREATION_USES_RENAMES = 1\n };\n \n+enum convert_aware_status {\n+\tCONVERT_AWARE_STATUS_NEVER = 0,\n+\tCONVERT_AWARE_STATUS_IN_PROCESS,\n+\tCONVERT_AWARE_STATUS_ALL\n+};\n+\n struct repo_config_values {\n \t/* section \"core\" config values */\n \tchar *attributes_file;\n@@ -130,6 +136,7 @@ struct repo_config_values {\n \tint apply_sparse_checkout;\n \tint trust_ctime;\n \tint check_stat;\n+\tenum convert_aware_status convert_aware_status;\n \tint zlib_compression_level;\n \tint pack_compression_level;\n \tint precomposed_unicode;\ndiff --git a/read-cache.c b/read-cache.c\nindex 8875706d8..2bb27a388 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -554,9 +554,21 @@ static int size_change_is_conclusive(struct index_state *istate,\n \t\t\t\t     const struct cache_entry *ce,\n \t\t\t\t     struct stat *st)\n {\n+\tstruct repo_config_values *cfg = repo_config_values(the_repository);\n+\tstruct conv_attrs ca;\n+\n+\tif (cfg->convert_aware_status == CONVERT_AWARE_STATUS_NEVER)\n+\t\treturn 1;\n+\n \tif (!S_ISREG(st->st_mode))\n \t\treturn 1;\n \n+\tif (cfg->convert_aware_status == CONVERT_AWARE_STATUS_IN_PROCESS) {\n+\t\tconvert_attrs(istate, &ca, ce->name);\n+\t\tif (ca.drv)\n+\t\t\treturn 1;\n+\t}\n+\n \treturn !would_convert_to_git(istate, ce->name);\n }\n \ndiff --git a/t/t0020-crlf.sh b/t/t0020-crlf.sh\nindex 88b728d50..4c127207e 100755\n--- a/t/t0020-crlf.sh\n+++ b/t/t0020-crlf.sh\n@@ -442,4 +442,51 @@ test_expect_success 'status sizes a text file by its CRLF pairs, not its CRs' '\n \t)\n '\n \n+test_expect_success 'core.convertAwareStatus=false restores the size shortcut' '\n+\tgit init convert-aware &&\n+\t(\n+\t\tcd convert-aware &&\n+\t\techo \"* text eol=lf\" >.gitattributes &&\n+\t\tprintf \"one\\ntwo\\nthree\\n\" >file.txt &&\n+\t\tgit add .gitattributes file.txt &&\n+\t\tgit commit -m initial &&\n+\t\tprintf \"one\\r\\ntwo\\r\\nthree\\r\\n\" >file.txt &&\n+\n+\t\tgit -c core.convertAwareStatus=false status --porcelain -uno >actual &&\n+\t\techo \" M file.txt\" >expect &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\tgit -c core.convertAwareStatus=true status --porcelain -uno >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n+\n+test_expect_success 'core.convertAwareStatus=no-filter leaves clean filters alone' '\n+\tgit init convert-aware-filter &&\n+\t(\n+\t\tcd convert-aware-filter &&\n+\t\twrite_script stripcr <<-\\EOF &&\n+\t\ttr -d \"\\015\"\n+\t\tEOF\n+\t\techo \"file.txt filter=stripcr\" >.gitattributes &&\n+\t\tgit config filter.stripcr.clean ./stripcr &&\n+\t\tprintf \"one\\ntwo\\nthree\\n\" >file.txt &&\n+\t\tgit add .gitattributes file.txt &&\n+\t\tgit commit -m initial &&\n+\t\tprintf \"one\\r\\ntwo\\r\\nthree\\r\\n\" >file.txt &&\n+\n+\t\tgit -c core.convertAwareStatus=no-filter status --porcelain -uno >actual &&\n+\t\techo \" M file.txt\" >expect &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\tgit status --porcelain -uno >actual &&\n+\t\ttest_must_be_empty actual\n+\t)\n+'\n+\n+test_expect_success 'core.convertAwareStatus rejects an unknown value' '\n+\ttest_must_fail git -c core.convertAwareStatus=bogus status 2>err &&\n+\ttest_grep \"invalid value\" err\n+'\n+\n test_done\n-- \n2.53.0\n\n\n"},{"id":"554551","messageId":"xmqqfqyfwi20.fsf@gitster.g","threadId":"66492","inReplyTo":"20261008204603.1988-2-curtis.allen.smith@gmail.com","subject":"Re: [PATCH 1/2] read-cache: do not trust a size change when conversion is active","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-10-09T05:37:27Z","receivedAt":"2026-10-09T05:37:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Curtis Allen Smith <curtis.allen.smith@gmail.com> writes:\n\n> \"git status\" can report a file as modified while \"git diff\" and\n> ...\n> next run of the tool flags everything again.\n>\n> Signed-off-by: Curtis Allen Smith <curtis.allen.smith@gmail.com>\n> ---\n\nThat's overly verbose.\n\n>  read-cache.c    | 129 ++++++++++++++++++++++++++++++++++++++++++++++--\n>  t/t0020-crlf.sh |  45 +++++++++++++++++\n>  2 files changed, 171 insertions(+), 3 deletions(-)\n\nAnd it is curious why we need so much new code, especially after\nreading an explaination in the proposed log message that makes it\nsound as if \"we let ce_modified_check_fs() to compare converted\nresult already when timestamps differ, and it is just the matter of\ndoing the same when sizes are the same\" is what is happening in the\npatch.  Why do we need to add a new function that compares converted\ndata?  A new function is not automatically a bad thing.  If there is\nalready an existing code path that does the same thing, a new\nfunction may be a good way to replace that code path with a more\ngeneric code and apply essentially the same logic implemented by\nthat new more generic code to a new code path.  But in such a\nrefactoring patch, we usually see a comparable number of removed\nlines, which is not what we see in the diffstat above.\n\n\n"}]}