# [RFC] Preserve per-file historical mtimes from Git history

2 messages from 2026-10-01 to 2026-10-02. Participants: Artem S. Tashkinov, Junio C Hamano.
Thread: https://gitlist.dev/t/66433

## Artem S. Tashkinov, 2026-10-01 09:43

Subject: [RFC] Preserve per-file historical mtimes from Git history
Message-ID: <011ad911-2d62-4179-b3a7-e14aab79f18a@gmx.com>

```
Hi,

I'd like to propose a way for Git to restore meaningful per-file
modification timestamps from repository history.

The problem is simple: after a clone, checkout, or creation of a source
archive, files normally receive timestamps related to the checkout or
archive operation rather than timestamps reflecting when their contents
last changed in Git.

For many repositories this does not matter much. For some, however, it
throws away useful information.

A particularly good example is linux-firmware. It contains thousands of
firmware blobs which have been independently added or updated over many
years. After obtaining a snapshot, ordinary filesystem tools no longer
tell me whether a particular blob was changed last week or has been
untouched for five years.

I would like Git to provide an optional operation along the lines of:

     git restore-mtimes

and possibly convenience options such as:

     git clone --historical-mtimes ...
     git checkout --historical-mtimes ...

The names are only examples; I am primarily interested in whether the
underlying behavior makes sense as a Git feature.

What I mean by a "historical mtime" is not an original filesystem mtime,
since Git obviously does not store those.

Instead, it would be a timestamp derived deterministically from Git
history:

   * find the newest commit in which the file's blob contents actually
     changed;
   * use that commit's timestamp as the file's mtime;
   * follow renames where Git can identify them;
   * do not change the timestamp for a pure rename;
   * do not change it for a mode-only change;
   * do change it for a rename accompanied by a content change.

There would also need to be a defined choice between author and
committer timestamps. I would personally expect the committer timestamp
to be the default, but I do not have a strong objection to making this
configurable.

For example, suppose a repository contains:

     firmware-a.bin   last content change: 2026-09-28
     firmware-b.bin   last content change: 2024-03-11
     firmware-c.bin   last content change: 2021-07-19

A normal checkout makes those historical differences invisible at the
filesystem level.

With historical mtimes restored, ordinary tools such as:

     ls -l
     stat
     find -newermt ...

would once again provide useful information.

There are also packaging and reproducible-build use cases. I originally
ran into this while looking at Fedora packages derived from Git
repositories. Source snapshot creation commonly destroys useful
per-file timestamp information, even though Git history contains enough
information to reconstruct a deterministic approximation.

Since the result depends only on repository history, two checkouts of
the same commit and history could receive the same mtimes. This is
different from simply preserving checkout time and seems useful for
reproducible source generation as well.

I wrote a proof-of-concept shell implementation which does roughly this.
It compares blob IDs rather than trusting status letters, ignores pure
renames and mode-only changes, handles symlinks without dereferencing
them, and attempts to give sensible behavior for merges.

The obvious disadvantage of doing this externally is efficiency. A
shell implementation tends to ask Git about file histories repeatedly.
Git itself should be able to compute the same information much more
efficiently in a single history traversal or using internal revision
machinery.

There are some semantics that I think deserve discussion before
implementation:

   1. Should author or committer timestamp be the default?

   2. What should the precise merge semantics be?

      My current interpretation is that a merge should count as a content
      change only if it introduces file contents not already present in
      the relevant parent history. A merge which merely selects an
      existing parent's blob should not make every such file appear newly
      modified.

   3. How should rename following behave across complicated/non-linear
      history?

      I realize Git does not store renames; they are inferred, so this
      can never be a perfect reconstruction.

   4. Should this be a standalone command, an option to checkout/restore,
      an archive feature, or some combination of these?

   5. Would it make sense for `git archive` to optionally use these
      historical per-file timestamps for archive members?

The last point may actually be more useful than clone-specific behavior.
For distributions and other source consumers, being able to do
something conceptually like:

     git archive --historical-mtimes <commit>

would preserve this information directly in generated source archives.

I am not suggesting that Git change its existing checkout behavior by
default. This would be opt-in, because some build systems intentionally
depend on freshly-created mtimes.

I mainly want to ask whether the Git project considers this information
useful enough to expose natively, and if so, what interface and 
semantics would fit Git best.

If there is interest, I can publish the proof-of-concept implementation
and its tests as a starting point for discussion.

Thanks,
Artem

```

## Junio C Hamano, 2026-10-02 01:28

Subject: Re: [RFC] Preserve per-file historical mtimes from Git history
Message-ID: <xmqqqzi83n86.fsf@gitster.g>
In-Reply-To: <011ad911-2d62-4179-b3a7-e14aab79f18a@gmx.com>

```
"Artem S. Tashkinov" <aros@gmx.com> writes:

Omitting parts and details that are not interesting (to me) from the
list.

>    3. How should rename following behave across complicated/non-linear
>       history?

You'd also want to think about what should happen when somebody
concatenates two files to create one.  No new contents are
introduced to the tree.

As to the definition of the historical times, I would find it a lot
more natural if they behaved as if you did this:

 * Prepare "git rev-list --first-parent --reverse HEAD" to prepare
   a list of commits that historically appeared at the tip of the
   trunk.

 * Make a checkout of the first commit in the list; give all the
   working tree files the committer timestamp of that commit.

 * Repeatedly run "git checkout" of the next commit, which should
   replace the paths that differ from the previous commit.  Give
   these paths the committer timestamp of the current commit.
   Repeat this process until you run out of the list.

and then looked at timestamp of each surviving paths in the working
tree.

>    5. Would it make sense for `git archive` to optionally use these
>       historical per-file timestamps for archive members?

Yes, it would be interesting to teach "git archive" to produce a
tarball with these synthesized timestamps.

```
