{"thread":{"id":"66433","subject":"[RFC] Preserve per-file historical mtimes from Git history","startedAt":"2026-10-01T09:43:28Z","lastAt":"2026-10-02T01:28:12Z","messageCount":2,"participants":["Artem S. Tashkinov","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"553821","messageId":"011ad911-2d62-4179-b3a7-e14aab79f18a@gmx.com","threadId":"66433","inReplyTo":null,"subject":"[RFC] Preserve per-file historical mtimes from Git history","fromName":"Artem S. Tashkinov","fromEmail":"aros@gmx.com","sentAt":"2026-10-01T09:43:25Z","receivedAt":"2026-10-01T09:43:28Z","isPatch":false,"body":"Hi,\n\nI'd like to propose a way for Git to restore meaningful per-file\nmodification timestamps from repository history.\n\nThe problem is simple: after a clone, checkout, or creation of a source\narchive, files normally receive timestamps related to the checkout or\narchive operation rather than timestamps reflecting when their contents\nlast changed in Git.\n\nFor many repositories this does not matter much. For some, however, it\nthrows away useful information.\n\nA particularly good example is linux-firmware. It contains thousands of\nfirmware blobs which have been independently added or updated over many\nyears. After obtaining a snapshot, ordinary filesystem tools no longer\ntell me whether a particular blob was changed last week or has been\nuntouched for five years.\n\nI would like Git to provide an optional operation along the lines of:\n\n     git restore-mtimes\n\nand possibly convenience options such as:\n\n     git clone --historical-mtimes ...\n     git checkout --historical-mtimes ...\n\nThe names are only examples; I am primarily interested in whether the\nunderlying behavior makes sense as a Git feature.\n\nWhat I mean by a \"historical mtime\" is not an original filesystem mtime,\nsince Git obviously does not store those.\n\nInstead, it would be a timestamp derived deterministically from Git\nhistory:\n\n   * find the newest commit in which the file's blob contents actually\n     changed;\n   * use that commit's timestamp as the file's mtime;\n   * follow renames where Git can identify them;\n   * do not change the timestamp for a pure rename;\n   * do not change it for a mode-only change;\n   * do change it for a rename accompanied by a content change.\n\nThere would also need to be a defined choice between author and\ncommitter timestamps. I would personally expect the committer timestamp\nto be the default, but I do not have a strong objection to making this\nconfigurable.\n\nFor example, suppose a repository contains:\n\n     firmware-a.bin   last content change: 2026-09-28\n     firmware-b.bin   last content change: 2024-03-11\n     firmware-c.bin   last content change: 2021-07-19\n\nA normal checkout makes those historical differences invisible at the\nfilesystem level.\n\nWith historical mtimes restored, ordinary tools such as:\n\n     ls -l\n     stat\n     find -newermt ...\n\nwould once again provide useful information.\n\nThere are also packaging and reproducible-build use cases. I originally\nran into this while looking at Fedora packages derived from Git\nrepositories. Source snapshot creation commonly destroys useful\nper-file timestamp information, even though Git history contains enough\ninformation to reconstruct a deterministic approximation.\n\nSince the result depends only on repository history, two checkouts of\nthe same commit and history could receive the same mtimes. This is\ndifferent from simply preserving checkout time and seems useful for\nreproducible source generation as well.\n\nI wrote a proof-of-concept shell implementation which does roughly this.\nIt compares blob IDs rather than trusting status letters, ignores pure\nrenames and mode-only changes, handles symlinks without dereferencing\nthem, and attempts to give sensible behavior for merges.\n\nThe obvious disadvantage of doing this externally is efficiency. A\nshell implementation tends to ask Git about file histories repeatedly.\nGit itself should be able to compute the same information much more\nefficiently in a single history traversal or using internal revision\nmachinery.\n\nThere are some semantics that I think deserve discussion before\nimplementation:\n\n   1. Should author or committer timestamp be the default?\n\n   2. What should the precise merge semantics be?\n\n      My current interpretation is that a merge should count as a content\n      change only if it introduces file contents not already present in\n      the relevant parent history. A merge which merely selects an\n      existing parent's blob should not make every such file appear newly\n      modified.\n\n   3. How should rename following behave across complicated/non-linear\n      history?\n\n      I realize Git does not store renames; they are inferred, so this\n      can never be a perfect reconstruction.\n\n   4. Should this be a standalone command, an option to checkout/restore,\n      an archive feature, or some combination of these?\n\n   5. Would it make sense for `git archive` to optionally use these\n      historical per-file timestamps for archive members?\n\nThe last point may actually be more useful than clone-specific behavior.\nFor distributions and other source consumers, being able to do\nsomething conceptually like:\n\n     git archive --historical-mtimes <commit>\n\nwould preserve this information directly in generated source archives.\n\nI am not suggesting that Git change its existing checkout behavior by\ndefault. This would be opt-in, because some build systems intentionally\ndepend on freshly-created mtimes.\n\nI mainly want to ask whether the Git project considers this information\nuseful enough to expose natively, and if so, what interface and \nsemantics would fit Git best.\n\nIf there is interest, I can publish the proof-of-concept implementation\nand its tests as a starting point for discussion.\n\nThanks,\nArtem\n"},{"id":"553881","messageId":"xmqqqzi83n86.fsf@gitster.g","threadId":"66433","inReplyTo":"011ad911-2d62-4179-b3a7-e14aab79f18a@gmx.com","subject":"Re: [RFC] Preserve per-file historical mtimes from Git history","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-10-02T01:28:09Z","receivedAt":"2026-10-02T01:28:12Z","isPatch":false,"body":"\"Artem S. Tashkinov\" <aros@gmx.com> writes:\n\nOmitting parts and details that are not interesting (to me) from the\nlist.\n\n>    3. How should rename following behave across complicated/non-linear\n>       history?\n\nYou'd also want to think about what should happen when somebody\nconcatenates two files to create one.  No new contents are\nintroduced to the tree.\n\nAs to the definition of the historical times, I would find it a lot\nmore natural if they behaved as if you did this:\n\n * Prepare \"git rev-list --first-parent --reverse HEAD\" to prepare\n   a list of commits that historically appeared at the tip of the\n   trunk.\n\n * Make a checkout of the first commit in the list; give all the\n   working tree files the committer timestamp of that commit.\n\n * Repeatedly run \"git checkout\" of the next commit, which should\n   replace the paths that differ from the previous commit.  Give\n   these paths the committer timestamp of the current commit.\n   Repeat this process until you run out of the list.\n\nand then looked at timestamp of each surviving paths in the working\ntree.\n\n>    5. Would it make sense for `git archive` to optionally use these\n>       historical per-file timestamps for archive members?\n\nYes, it would be interesting to teach \"git archive\" to produce a\ntarball with these synthesized timestamps.\n"}]}