Volume XXII, number 279Tuesday, October 6, 2026Latest message 33 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

[RFC] Preserve per-file historical mtimes from Git history

2 messages between Oct 1, 2026 and Oct 2, 2026, from Artem S. Tashkinov, Junio C Hamano.

Plain Markdown or JSON for tools and agents.

Artem S. TashkinovOct 1, 2026, 09:43 UTC on lore
Hi,

I'd like to propose a way for Git to restore meaningful per-file modification timestamps from repository history.

The problem is simple: after a clone, checkout, or creation of a source archive, files normally receive timestamps related to the checkout or archive operation rather than timestamps reflecting when their contents last changed in Git.

For many repositories this does not matter much. For some, however, it throws away useful information.

A particularly good example is linux-firmware. It contains thousands of firmware blobs which have been independently added or updated over many years. After obtaining a snapshot, ordinary filesystem tools no longer tell me whether a particular blob was changed last week or has been untouched for five years.

I would like Git to provide an optional operation along the lines of:
     git restore-mtimes
and possibly convenience options such as:
     git clone --historical-mtimes ...
     git checkout --historical-mtimes ...

The names are only examples; I am primarily interested in whether the underlying behavior makes sense as a Git feature.

What I mean by a "historical mtime" is not an original filesystem mtime, since Git obviously does not store those.

Instead, it would be a timestamp derived deterministically from Git history:

   * find the newest commit in which the file's blob contents actually
     changed;
   * use that commit's timestamp as the file's mtime;
   * follow renames where Git can identify them;
   * do not change the timestamp for a pure rename;
   * do not change it for a mode-only change;
   * do change it for a rename accompanied by a content change.

There would also need to be a defined choice between author and committer timestamps. I would personally expect the committer timestamp to be the default, but I do not have a strong objection to making this configurable.

For example, suppose a repository contains:
     firmware-a.bin   last content change: 2026-09-28
     firmware-b.bin   last content change: 2024-03-11
     firmware-c.bin   last content change: 2021-07-19

A normal checkout makes those historical differences invisible at the filesystem level.

With historical mtimes restored, ordinary tools such as:
     ls -l
     stat
     find -newermt ...
would once again provide useful information.

There are also packaging and reproducible-build use cases. I originally ran into this while looking at Fedora packages derived from Git repositories. Source snapshot creation commonly destroys useful per-file timestamp information, even though Git history contains enough information to reconstruct a deterministic approximation.

Since the result depends only on repository history, two checkouts of the same commit and history could receive the same mtimes. This is different from simply preserving checkout time and seems useful for reproducible source generation as well.

I wrote a proof-of-concept shell implementation which does roughly this. It compares blob IDs rather than trusting status letters, ignores pure renames and mode-only changes, handles symlinks without dereferencing them, and attempts to give sensible behavior for merges.

The obvious disadvantage of doing this externally is efficiency. A shell implementation tends to ask Git about file histories repeatedly. Git itself should be able to compute the same information much more efficiently in a single history traversal or using internal revision machinery.

There are some semantics that I think deserve discussion before implementation:

   1. Should author or committer timestamp be the default?
   2. What should the precise merge semantics be?
      My current interpretation is that a merge should count as a content
      change only if it introduces file contents not already present in
      the relevant parent history. A merge which merely selects an
      existing parent's blob should not make every such file appear newly
      modified.
   3. How should rename following behave across complicated/non-linear
      history?
      I realize Git does not store renames; they are inferred, so this
      can never be a perfect reconstruction.
   4. Should this be a standalone command, an option to checkout/restore,
      an archive feature, or some combination of these?
   5. Would it make sense for `git archive` to optionally use these
      historical per-file timestamps for archive members?

The last point may actually be more useful than clone-specific behavior. For distributions and other source consumers, being able to do something conceptually like:

     git archive --historical-mtimes <commit>
would preserve this information directly in generated source archives.

I am not suggesting that Git change its existing checkout behavior by default. This would be opt-in, because some build systems intentionally depend on freshly-created mtimes.

I mainly want to ask whether the Git project considers this information useful enough to expose natively, and if so, what interface and semantics would fit Git best.

If there is interest, I can publish the proof-of-concept implementation and its tests as a starting point for discussion.

Thanks, Artem

Junio C HamanoOct 2, 2026, 01:28 UTC in reply to Artem S. Tashkinov on lore

Re: [RFC] Preserve per-file historical mtimes from Git history

"Artem S. Tashkinov" <aros@gmx.com> writes:

Omitting parts and details that are not interesting (to me) from the list.

>    3. How should rename following behave across complicated/non-linear
>       history?

You'd also want to think about what should happen when somebody concatenates two files to create one. No new contents are introduced to the tree.

As to the definition of the historical times, I would find it a lot more natural if they behaved as if you did this:

 * Prepare "git rev-list --first-parent --reverse HEAD" to prepare
   a list of commits that historically appeared at the tip of the
   trunk.
 * Make a checkout of the first commit in the list; give all the
   working tree files the committer timestamp of that commit.
 * Repeatedly run "git checkout" of the next commit, which should
   replace the paths that differ from the previous commit.  Give
   these paths the committer timestamp of the current commit.
   Repeat this process until you run out of the list.

and then looked at timestamp of each surviving paths in the working tree.

>    5. Would it make sense for `git archive` to optionally use these
>       historical per-file timestamps for archive members?

Yes, it would be interesting to teach "git archive" to produce a tarball with these synthesized timestamps.

Back to recent threads