Hi,
I'd like to propose a way for Git to restore meaningful per-file modification timestamps from repository history.
The problem is simple: after a clone, checkout, or creation of a source archive, files normally receive timestamps related to the checkout or archive operation rather than timestamps reflecting when their contents last changed in Git.
For many repositories this does not matter much. For some, however, it throws away useful information.
A particularly good example is linux-firmware. It contains thousands of firmware blobs which have been independently added or updated over many years. After obtaining a snapshot, ordinary filesystem tools no longer tell me whether a particular blob was changed last week or has been untouched for five years.
I would like Git to provide an optional operation along the lines of:
git restore-mtimes
and possibly convenience options such as:
git clone --historical-mtimes ...
git checkout --historical-mtimes ...The names are only examples; I am primarily interested in whether the underlying behavior makes sense as a Git feature.
What I mean by a "historical mtime" is not an original filesystem mtime, since Git obviously does not store those.
Instead, it would be a timestamp derived deterministically from Git history:
* find the newest commit in which the file's blob contents actually
changed;
* use that commit's timestamp as the file's mtime;
* follow renames where Git can identify them;
* do not change the timestamp for a pure rename;
* do not change it for a mode-only change;
* do change it for a rename accompanied by a content change.There would also need to be a defined choice between author and committer timestamps. I would personally expect the committer timestamp to be the default, but I do not have a strong objection to making this configurable.
For example, suppose a repository contains:
firmware-a.bin last content change: 2026-09-28
firmware-b.bin last content change: 2024-03-11
firmware-c.bin last content change: 2021-07-19A normal checkout makes those historical differences invisible at the filesystem level.
With historical mtimes restored, ordinary tools such as:
ls -l
stat
find -newermt ...would once again provide useful information.
There are also packaging and reproducible-build use cases. I originally ran into this while looking at Fedora packages derived from Git repositories. Source snapshot creation commonly destroys useful per-file timestamp information, even though Git history contains enough information to reconstruct a deterministic approximation.
Since the result depends only on repository history, two checkouts of the same commit and history could receive the same mtimes. This is different from simply preserving checkout time and seems useful for reproducible source generation as well.
I wrote a proof-of-concept shell implementation which does roughly this. It compares blob IDs rather than trusting status letters, ignores pure renames and mode-only changes, handles symlinks without dereferencing them, and attempts to give sensible behavior for merges.
The obvious disadvantage of doing this externally is efficiency. A shell implementation tends to ask Git about file histories repeatedly. Git itself should be able to compute the same information much more efficiently in a single history traversal or using internal revision machinery.
There are some semantics that I think deserve discussion before implementation:
1. Should author or committer timestamp be the default?
2. What should the precise merge semantics be?
My current interpretation is that a merge should count as a content
change only if it introduces file contents not already present in
the relevant parent history. A merge which merely selects an
existing parent's blob should not make every such file appear newly
modified. 3. How should rename following behave across complicated/non-linear
history? I realize Git does not store renames; they are inferred, so this
can never be a perfect reconstruction. 4. Should this be a standalone command, an option to checkout/restore,
an archive feature, or some combination of these? 5. Would it make sense for `git archive` to optionally use these
historical per-file timestamps for archive members?The last point may actually be more useful than clone-specific behavior. For distributions and other source consumers, being able to do something conceptually like:
git archive --historical-mtimes <commit>
would preserve this information directly in generated source archives.
I am not suggesting that Git change its existing checkout behavior by default. This would be opt-in, because some build systems intentionally depend on freshly-created mtimes.
I mainly want to ask whether the Git project considers this information useful enough to expose natively, and if so, what interface and semantics would fit Git best.
If there is interest, I can publish the proof-of-concept implementation and its tests as a starting point for discussion.
Thanks, Artem