{"thread":{"id":"66283","subject":"[PATCH] git pull silently overwrites local directory with symlink due to .gitignore \"dir/\"","startedAt":"2026-09-07T08:19:19Z","lastAt":"2026-09-07T19:41:10Z","messageCount":2,"participants":["AIKSXD ax","brian m. carlson"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"552103","messageId":"DSWPR04MB9945756976C15A3A4978CE9AD0B22@DSWPR04MB9945.namprd04.prod.outlook.com","threadId":"66283","inReplyTo":null,"subject":"[PATCH] git pull silently overwrites local directory with symlink due to .gitignore \"dir/\"","fromName":"AIKSXD ax","fromEmail":"aiksxd@outlook.com","sentAt":"2026-09-07T08:19:17Z","receivedAt":"2026-09-07T08:19:19Z","isPatch":true,"body":"Hello,I would like to report a issue in Git that can cause silent data loss on user machines. The problem occurs when a '.gitignore' pattern ending with a slash (e.g. 'dir/') is used to ignore a directory, but a symbolic link with the same name will be committed. Later, when another user pulls the repository, Git silently replaces their local directory with that symlink, destroying all data inside it without any hints.\n\nOS: Linux(Git 2.43.0) & Windows(Git 2.53.0.windows.2) both reproduced\n\nConcrete example (from a real incident):\n1. We had a repository with a symlink named 'dataset' pointing to a large data directory located outside the repo:\n\n ➜  experiment git:(main) ✗ ll\n        total 0\n        lrwxrwxrwx 1 ax ax 10 Sep  7 14:11 dataset -> ../dataset\n        -rw-r--r-- 1 ax ax  0 Sep  7 14:17 train.py\n\n➜  experiment git:(main) ✗ cat .gitignore\n        dataset/\n\n➜  experiment git:(main) ✗ git add . && git commit -m \"feat: ...\" && git push\n        [main 6a9dca8] feat: ...\n        1 file changed, 1 insertion(+)\n        create mode 120000 dataset\n\n-------------\n2. On another machine, the same repository had a real directory named 'dataset' containing important data. After `git pull`, Git replaced that directory with the symlink without any warning:\n\n➜  experiment git:(main) du -h -d 0 dataset\n        64M     dataset\n\n➜  experiment git:(main) ll\n        total 4.0K\n        drwxr-xr-x 3 ax ax 4.0K Sep  7 14:32 dataset\n        -rw-r--r-- 1 ax ax    0 Sep  7 14:31 train.py\n\n➜  experiment git:(main) git pull\n        remote: Enumerating objects: 4, done.\n        remote: Counting objects: 100% (4/4), done.\n        remote: Compressing objects: 100% (2/2), done.\n        remote: Total 3 (delta 0), reused 3 (delta 0), pack-reused 0 (from 0)\n        Unpacking objects: 100% (3/3), 292 bytes | 292.00 KiB/s, done.\n        From github.com:aiksxd/experiment\n           0bdcba7..6a9dca8  main       -> origin/main\n        Updating 0bdcba7..6a9dca8\n        Fast-forward\n         dataset | 1 +\n         1 file changed, 1 insertion(+)\n         create mode 120000 dataset\n\n➜  experiment git:(main) ll\n        total 0\n        lrwxrwxrwx 1 ax ax 10 Sep  7 14:38 dataset -> ../dataset\n        -rw-r--r-- 1 ax ax  0 Sep  7 14:31 train.py\n\n➜  experiment git:(main) du -h -d 0 dataset\n        0       dataset\n\n-----------\n- The symlink is tracked and committed because the trailing-slash ignore rule have no effect on files.\n- On pull, Git silently replaces the local directory with the symlink, causing irreversible data loss.\nThis is unacceptable behavior; Git should never overwrite a local directory with a symlink without explicit user confirmation.\n\nImpact:\nThis issue can result in the loss of hundreds of gigabytes of local data, as users often keep large datasets or other important directories with the same name as an ignored symlink. The data loss is silent and occurs during a routine 'git pull'( I don’t know why so much free space showed up on my computer that day).\n\nMy options:\nThe pattern 'dataset/' should also ignore a symlink with that name, so it never enters the repository in the first place.\nIf such a symlink is committed (accidentally or otherwise), Git must detect the conflict when pulling to a machine that has a real directory at the same path, and refuse to overwrite it without prompting.\n\nThank you for your time and for maintaining Git.\n\nBest regards,\naiksxd@126.com"},{"id":"552156","messageId":"ap8TTRctdrsFo2l1@fruit.crustytoothpaste.net","threadId":"66283","inReplyTo":"DSWPR04MB9945756976C15A3A4978CE9AD0B22@DSWPR04MB9945.namprd04.prod.outlook.com","subject":"Re: [PATCH] git pull silently overwrites local directory with symlink due to .gitignore \"dir/\"","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-09-07T19:41:01Z","receivedAt":"2026-09-07T19:41:10Z","isPatch":true,"body":"On 2026-09-07 at 08:19:17, AIKSXD ax wrote:\n> Hello,I would like to report a issue in Git that can cause silent data loss on user machines. The problem occurs when a '.gitignore' pattern ending with a slash (e.g. 'dir/') is used to ignore a directory, but a symbolic link with the same name will be committed. Later, when another user pulls the repository, Git silently replaces their local directory with that symlink, destroying all data inside it without any hints.\n> \n> OS: Linux(Git 2.43.0) & Windows(Git 2.53.0.windows.2) both reproduced\n> - The symlink is tracked and committed because the trailing-slash ignore rule have no effect on files.\n\nYes, as you've noticed, symlinks (and files) are not ignored by patterns\ncontaining a trailing slash.  This is because `dataset/` doesn't\nactually ignore `dataset`, but everything under it instead.  The index\ndoesn't track directories, only regular files and symlinks, so `dataset`\nas a symlink is not even considered by that rule.\n\n> - On pull, Git silently replaces the local directory with the symlink, causing irreversible data loss.\n> This is unacceptable behavior; Git should never overwrite a local directory with a symlink without explicit user confirmation.\n\nI tested this with a non-symlink file and Git also removes the directory\nin this case.  As you noticed, `git checkout` deletes ignored files and\ndirectories.  You can see in the manual page:\n\n     --overwrite-ignore, --no-overwrite-ignore\n         Silently overwrite ignored files when switching branches. This\n         is the default behavior. Use --no-overwrite-ignore to abort the\n         operation when the new branch contains ignored files.\n\nGit does not consider ignored files to be valuable by default.  There\nhas been discussion of a `precious` attribute to preserve ignored files,\nbut it hasn't been implemented yet.  [0] is one relatively recent proposal.\n\n> Impact:\n> This issue can result in the loss of hundreds of gigabytes of local data, as users often keep large datasets or other important directories with the same name as an ignored symlink. The data loss is silent and occurs during a routine 'git pull'( I don’t know why so much free space showed up on my computer that day).\n> \n> My options:\n> The pattern 'dataset/' should also ignore a symlink with that name, so it never enters the repository in the first place.\n\nThis would be a substantial change in behaviour.  We don't know that the\nsymlink points to a directory and it would be bizarre to have a symlink\nto a file affected in that way.  Moreover, on Unix, a symlink need not\nactually point _anywhere_, so whether this worked would be dependent on\nsubtle behaviour about what the symlink is pointing to at this time.\n\n> If such a symlink is committed (accidentally or otherwise), Git must detect the conflict when pulling to a machine that has a real directory at the same path, and refuse to overwrite it without prompting.\n\nThis would also be a big change in behaviour and probably break a lot of\ntooling that relies on the status quo.\n\nIn general, I would recommend not keeping valuable data in untracked or\nignored files within the working tree.  I've seen lots of data loss from\nthis case and have actually had to restore people's development VMs from\na snapshot for that reason at a previous job.\n\n[0] https://lore.kernel.org/git/pull.1627.git.1703643931314.gitgitgadget@gmail.com/\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"}]}