# [PATCH] git pull silently overwrites local directory with symlink due to .gitignore "dir/"

2 messages from 2026-09-07 to 2026-09-07. Participants: AIKSXD ax, brian m. carlson.
Thread: https://gitlist.dev/t/66283

## AIKSXD ax, 2026-09-07 08:19

Subject: [PATCH] git pull silently overwrites local directory with symlink due to .gitignore "dir/"
Message-ID: <DSWPR04MB9945756976C15A3A4978CE9AD0B22@DSWPR04MB9945.namprd04.prod.outlook.com>

```
Hello,I would like to report a issue in Git that can cause silent data loss on user machines. The problem occurs when a '.gitignore' pattern ending with a slash (e.g. 'dir/') is used to ignore a directory, but a symbolic link with the same name will be committed. Later, when another user pulls the repository, Git silently replaces their local directory with that symlink, destroying all data inside it without any hints.

OS: Linux(Git 2.43.0) & Windows(Git 2.53.0.windows.2) both reproduced

Concrete example (from a real incident):
1. We had a repository with a symlink named 'dataset' pointing to a large data directory located outside the repo:

 ➜  experiment git:(main) ✗ ll
        total 0
        lrwxrwxrwx 1 ax ax 10 Sep  7 14:11 dataset -> ../dataset
        -rw-r--r-- 1 ax ax  0 Sep  7 14:17 train.py

➜  experiment git:(main) ✗ cat .gitignore
        dataset/

➜  experiment git:(main) ✗ git add . && git commit -m "feat: ..." && git push
        [main 6a9dca8] feat: ...
        1 file changed, 1 insertion(+)
        create mode 120000 dataset

-------------
2. On another machine, the same repository had a real directory named 'dataset' containing important data. After `git pull`, Git replaced that directory with the symlink without any warning:

➜  experiment git:(main) du -h -d 0 dataset
        64M     dataset

➜  experiment git:(main) ll
        total 4.0K
        drwxr-xr-x 3 ax ax 4.0K Sep  7 14:32 dataset
        -rw-r--r-- 1 ax ax    0 Sep  7 14:31 train.py

➜  experiment git:(main) git pull
        remote: Enumerating objects: 4, done.
        remote: Counting objects: 100% (4/4), done.
        remote: Compressing objects: 100% (2/2), done.
        remote: Total 3 (delta 0), reused 3 (delta 0), pack-reused 0 (from 0)
        Unpacking objects: 100% (3/3), 292 bytes | 292.00 KiB/s, done.
        From github.com:aiksxd/experiment
           0bdcba7..6a9dca8  main       -> origin/main
        Updating 0bdcba7..6a9dca8
        Fast-forward
         dataset | 1 +
         1 file changed, 1 insertion(+)
         create mode 120000 dataset

➜  experiment git:(main) ll
        total 0
        lrwxrwxrwx 1 ax ax 10 Sep  7 14:38 dataset -> ../dataset
        -rw-r--r-- 1 ax ax  0 Sep  7 14:31 train.py

➜  experiment git:(main) du -h -d 0 dataset
        0       dataset

-----------
- The symlink is tracked and committed because the trailing-slash ignore rule have no effect on files.
- On pull, Git silently replaces the local directory with the symlink, causing irreversible data loss.
This is unacceptable behavior; Git should never overwrite a local directory with a symlink without explicit user confirmation.

Impact:
This issue can result in the loss of hundreds of gigabytes of local data, as users often keep large datasets or other important directories with the same name as an ignored symlink. The data loss is silent and occurs during a routine 'git pull'( I don’t know why so much free space showed up on my computer that day).

My options:
The pattern 'dataset/' should also ignore a symlink with that name, so it never enters the repository in the first place.
If such a symlink is committed (accidentally or otherwise), Git must detect the conflict when pulling to a machine that has a real directory at the same path, and refuse to overwrite it without prompting.

Thank you for your time and for maintaining Git.

Best regards,
aiksxd@126.com
```

## brian m. carlson, 2026-09-07 19:41

Subject: Re: [PATCH] git pull silently overwrites local directory with symlink due to .gitignore "dir/"
Message-ID: <ap8TTRctdrsFo2l1@fruit.crustytoothpaste.net>
In-Reply-To: <DSWPR04MB9945756976C15A3A4978CE9AD0B22@DSWPR04MB9945.namprd04.prod.outlook.com>

```
On 2026-09-07 at 08:19:17, AIKSXD ax wrote:
> Hello,I would like to report a issue in Git that can cause silent data loss on user machines. The problem occurs when a '.gitignore' pattern ending with a slash (e.g. 'dir/') is used to ignore a directory, but a symbolic link with the same name will be committed. Later, when another user pulls the repository, Git silently replaces their local directory with that symlink, destroying all data inside it without any hints.
> 
> OS: Linux(Git 2.43.0) & Windows(Git 2.53.0.windows.2) both reproduced
> - The symlink is tracked and committed because the trailing-slash ignore rule have no effect on files.

Yes, as you've noticed, symlinks (and files) are not ignored by patterns
containing a trailing slash.  This is because `dataset/` doesn't
actually ignore `dataset`, but everything under it instead.  The index
doesn't track directories, only regular files and symlinks, so `dataset`
as a symlink is not even considered by that rule.

> - On pull, Git silently replaces the local directory with the symlink, causing irreversible data loss.
> This is unacceptable behavior; Git should never overwrite a local directory with a symlink without explicit user confirmation.

I tested this with a non-symlink file and Git also removes the directory
in this case.  As you noticed, `git checkout` deletes ignored files and
directories.  You can see in the manual page:

     --overwrite-ignore, --no-overwrite-ignore
         Silently overwrite ignored files when switching branches. This
         is the default behavior. Use --no-overwrite-ignore to abort the
         operation when the new branch contains ignored files.

Git does not consider ignored files to be valuable by default.  There
has been discussion of a `precious` attribute to preserve ignored files,
but it hasn't been implemented yet.  [0] is one relatively recent proposal.

> Impact:
> This issue can result in the loss of hundreds of gigabytes of local data, as users often keep large datasets or other important directories with the same name as an ignored symlink. The data loss is silent and occurs during a routine 'git pull'( I don’t know why so much free space showed up on my computer that day).
> 
> My options:
> The pattern 'dataset/' should also ignore a symlink with that name, so it never enters the repository in the first place.

This would be a substantial change in behaviour.  We don't know that the
symlink points to a directory and it would be bizarre to have a symlink
to a file affected in that way.  Moreover, on Unix, a symlink need not
actually point _anywhere_, so whether this worked would be dependent on
subtle behaviour about what the symlink is pointing to at this time.

> If such a symlink is committed (accidentally or otherwise), Git must detect the conflict when pulling to a machine that has a real directory at the same path, and refuse to overwrite it without prompting.

This would also be a big change in behaviour and probably break a lot of
tooling that relies on the status quo.

In general, I would recommend not keeping valuable data in untracked or
ignored files within the working tree.  I've seen lots of data loss from
this case and have actually had to restore people's development VMs from
a snapshot for that reason at a previous job.

[0] https://lore.kernel.org/git/pull.1627.git.1703643931314.gitgitgadget@gmail.com/
-- 
brian m. carlson (they/them)
Toronto, Ontario, CA

```
