{"thread":{"id":"54156","subject":"Fastest way to set files date and time to latest commit time of each one","startedAt":"2020-08-29T01:36:26Z","lastAt":"2020-09-02T19:28:48Z","messageCount":6,"participants":["Ivan Baldo","Junio C Hamano","Eric Wong","Raymond E. Pasco","Andreas Schwab"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"404742","messageId":"CAEbcw=3mOoYuJo2mQgqB2aJgn-D2i_7ZRmhfPvYNVHD1Kp8wuA@mail.gmail.com","threadId":"54156","inReplyTo":null,"subject":"Fastest way to set files date and time to latest commit time of each one","fromName":"Ivan Baldo","fromEmail":"ibaldo@gmail.com","sentAt":"2020-08-29T01:36:10Z","receivedAt":"2020-08-29T01:36:26Z","isPatch":false,"sender":{"key":"ibaldo@gmail.com","avatar":null},"body":"  Hello.\n  I know this is not standard usage of git, but I need a way to have\nmore stable dates and times in the files in order to avoid rsync\nchecksumming.\n  So I found this\nhttps://stackoverflow.com/questions/2179722/checking-out-old-file-with-original-create-modified-timestamps/2179876#2179876\nand modified it a bit to run in CentOS 7:\n\nIFS=\"\n\"\nfor FILE in $(git ls-files -z | tr '\\0' '\\n')\ndo\n    TIME=$(git log --pretty=format:%cd -n 1 --date=iso -- \"$FILE\")\n    touch -c -m -d \"$TIME\" \"$FILE\"\ndone\n\n  Unfortunately it takes ages for a 84k files repo.\n  I see the CPU usage is dominated by the git log command.\n  I know a way I could use to split the work for all the CPU threads\nbut anyway, I would like to know if you guys and girls know of a\nfaster way to do this.\n  Also I know of other utilities that store the metadata in Git, but I\nam trying to avoid that for the moment.\n  Thanks a lot in advance!\n  Have a nice day.\nP.s.: please Cc replies to me.\n\n-- \nIvan Baldo - ibaldo@gmail.com - http://ibaldo.codigolibre.net/\nFreelance C++/PHP programmer and GNU/Linux systems administrator.\nThe sky isn't the limit!\n"},{"id":"404743","messageId":"xmqq8sdym93d.fsf@gitster.c.googlers.com","threadId":"54156","inReplyTo":"CAEbcw=3mOoYuJo2mQgqB2aJgn-D2i_7ZRmhfPvYNVHD1Kp8wuA@mail.gmail.com","subject":"Re: Fastest way to set files date and time to latest commit time of each one","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-08-29T03:20:22Z","receivedAt":"2020-08-29T03:20:29Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ivan Baldo <ibaldo@gmail.com> writes:\n\n>   I know this is not standard usage of git, but I need a way to have\n> more stable dates and times in the files in order to avoid rsync\n> checksumming.\n\nWould you care to elaborate a bit more about the use case?  From\nwhat you wrote, I would assume:\n\n - The source of the rsync transfer is a git working tree.  It often\n   has the checkout of the latest and greatest version, but during\n   development, it may switch to older commit (e.g. to find where\n   regression occurred) or not-yet-ready commit (e.g. work in\n   progress that is not given to upstream).  You check out the\n   version you want to sync to the destination before initiating\n   rsync.\n\n - The destination of the rsync transfer is meant to serve as a\n   back-up of the latest and greatest, periodical snapshot of a\n   branch, etc., which is NOT controlled by git and transfer does\n   not happen in the reverse direction [*2*]\n\nBecause the working tree of the source repository is used to check\nout different versions between rsync sessions, files that did not\nchange between the commit you sync'ed to destination the last time\nand the commit you are about to sync still may have been touched and\nhave different timestamp, requiring rsync to check the contents.\n\nAnd as a workaround, you are willing to change the workflow to\n\"touch\" the working tree files, immediately before you run the next\nrsync, in a predictable way so that the timestamp of a file whose\ncontents did not change since the last rsync session would have the\nsame timestamp.  This may break your build next time you run \"make\"\nin the source working tree (because your object files that are\nexcluded from your rsync may have newer timestamp than the\ncorresponding source even when they must be recompiled due to your\n\"touch\"ing), but you are willing to pay the cost of say \"make clean\"\nafter \"touch\"ing.\n\nIs that the kind of use case you have around \"rsync\"?\n\nTo the question \"what is the time this file was last modified?\",\nthere is no simple and cheap answer that is easy to explain to\nend-users, unless your development is completely linear [*2*].\n\nThe loop you showed would be the right one in a linear history, and\nwith recent development to record which paths were changed in each\ncommit in the commit-graph data structure, the script should work a\nlot faster than traditional git.\n\n\n[Footnote]\n\n*1* Otherwise, you'd be just mirror-fetching from the source\n    repository.  If that can be arranged, running \"git pull --ff-only\"\n    on the destination side to update from the source side would\n    be a lot more efficient than running rsync, I would imagine.\n\n\n*2* In a history with merges, because two or more branches can touch\n    the file in parallel development at different times and then the\n    resulting parallel histories get merged into a single history.\n    When two or more of these parallel history gave the file in\n    question an identical content at different times, and the merge\n    result was recorded as the same content, you'd need to follow\n    ALL the paths and compare the timestamp of these commits to pick\n    one (which one?  the oldest one?  the newest one?  does the\n    order of parents in the merge matter?)\n\n"},{"id":"404745","messageId":"20200829044842.GA5732@dcvr","threadId":"54156","inReplyTo":"CAEbcw=3mOoYuJo2mQgqB2aJgn-D2i_7ZRmhfPvYNVHD1Kp8wuA@mail.gmail.com","subject":"Re: Fastest way to set files date and time to latest commit time of each one","fromName":"Eric Wong","fromEmail":"e@yhbt.net","sentAt":"2020-08-29T04:48:42Z","receivedAt":"2020-08-29T05:03:00Z","isPatch":false,"sender":{"key":"e@yhbt.net","avatar":null},"body":"Ivan Baldo <ibaldo@gmail.com> wrote:\n>   Hello.\n>   I know this is not standard usage of git, but I need a way to have\n> more stable dates and times in the files in order to avoid rsync\n> checksumming.\n>   So I found this\n> https://stackoverflow.com/questions/2179722/checking-out-old-file-with-original-create-modified-timestamps/2179876#2179876\n> and modified it a bit to run in CentOS 7:\n> \n> IFS=\"\n> \"\n> for FILE in $(git ls-files -z | tr '\\0' '\\n')\n> do\n>     TIME=$(git log --pretty=format:%cd -n 1 --date=iso -- \"$FILE\")\n>     touch -c -m -d \"$TIME\" \"$FILE\"\n> done\n> \n>   Unfortunately it takes ages for a 84k files repo.\n>   I see the CPU usage is dominated by the git log command.\n\nrunning git log for each file isn't necessary.\n\nOn Debian, rsync actually ships the `git-set-file-times' script\nin /usr/share/doc/rsync/scripts/ which only runs `git log' once\nand parses it.\n\nYou can also get my (original) version from:\nhttps://yhbt.net/git-set-file-times\n\n>   I know a way I could use to split the work for all the CPU threads\n> but anyway, I would like to know if you guys and girls know of a\n> faster way to do this.\n\nMuch of your overhead is going to be from process spawning.\nMy Perl version reduces that significantly.\n\nI haven't tried it with 84K files, but it'll have to keep all\nthose filenames in memory.  I'm not sure if parallelizing\nutime() syscalls is worth it, either; maybe it helps on SSD\nmore than HDD.\n"},{"id":"404746","messageId":"C597RQ1V1DCG.3EEXY0665KK35@ziyou.local","threadId":"54156","inReplyTo":"xmqq8sdym93d.fsf@gitster.c.googlers.com","subject":"Re: Fastest way to set files date and time to latest commit time of each one","fromName":"Raymond E. Pasco","fromEmail":"ray@ameretat.dev","sentAt":"2020-08-29T04:59:58Z","receivedAt":"2020-08-29T05:03:00Z","isPatch":false,"sender":{"key":"ray@ameretat.dev","avatar":"https://avatars.githubusercontent.com/u/115765?v=4"},"body":"On Fri Aug 28, 2020 at 11:20 PM EDT, Junio C Hamano wrote:\n> - The source of the rsync transfer is a git working tree. It often\n> has the checkout of the latest and greatest version, but during\n> development, it may switch to older commit (e.g. to find where\n> regression occurred) or not-yet-ready commit (e.g. work in\n> progress that is not given to upstream). You check out the\n> version you want to sync to the destination before initiating\n> rsync.\n\nAssuming this is the case, perhaps a separate worktree that you touch\nless often being the source for rsync might save rsync some\nbandwidth/cpu checking things.\n"},{"id":"404747","messageId":"87bliu9cfd.fsf@linux-m68k.org","threadId":"54156","inReplyTo":"CAEbcw=3mOoYuJo2mQgqB2aJgn-D2i_7ZRmhfPvYNVHD1Kp8wuA@mail.gmail.com","subject":"Re: Fastest way to set files date and time to latest commit time of each one","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2020-08-29T06:46:46Z","receivedAt":"2020-08-29T06:46:53Z","isPatch":false,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"I'm using this script:\n\n#!/bin/sh\ngit log --name-only --format=format:%n%ct -- \"$@\" |\nperl -e 'my $do_date = 0; chomp(my $cdup = `git rev-parse --show-cdup`);\n    while (<>) {\n\tchomp;\n\tif ($do_date) {\n\t    next if ($_ eq \"\");\n\t    die \"Unexpected $_\\n\" unless /^[0-9]+$/;\n\t    $d = $_;\n\t    $do_date = 0;\n\t} elsif ($_ eq \"\") {\n\t    $do_date = 1;\n\t} elsif (!defined($seen{$_})) {\n\t    $seen{$_} = 1;\n \t    utime $d, $d, \"$cdup$_\";\n \t}\n    }'\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 7578 EB47 D4E5 4D69 2510  2552 DF73 E780 A9DA AEC1\n\"And now for something completely different.\"\n"},{"id":"404931","messageId":"CAEbcw=1S8CE-fnM91F0_KKCXqDzDy5B1Bycz2x2ymtZmi+miWw@mail.gmail.com","threadId":"54156","inReplyTo":"20200829044842.GA5732@dcvr","subject":"Re: Fastest way to set files date and time to latest commit time of each one","fromName":"Ivan Baldo","fromEmail":"ibaldo@gmail.com","sentAt":"2020-09-02T19:28:34Z","receivedAt":"2020-09-02T19:28:48Z","isPatch":false,"sender":{"key":"ibaldo@gmail.com","avatar":null},"body":"  Hello everyone!\n  I just now managed to get time to work again on this, sorry for\nreplying so late but wanted to do a single reply with the conclusion\nif possible.\n  Let me tell you that I feel very humbled by all your replies, thanks\na lot for your time and concern with my inquiry!\n  Eric's script is not only in Debian but also in CentOS 7 (and I\nguess Red Hat 7 too) in\n/usr/share/doc/rsync-*/support/git-set-file-times.\n  My use case is similar to his: a cluster of identically configured\nweb servers with autoscaling (tested up to 100 servers) which when\nthey boot (or there is a new version of any of the websites), rsync\nthe current version from another server.\n  So currently if we build the system image of the web servers in\ntandem with the central server everything works smoothly, the problem\nis when we recreate from scratch any of the pre-saved images, in which\ncase we get the dates mismatch and unnecessary rsync checksumming when\nput to production.\n  Will use Eric's script from CentOS 7 as-is from now on, to avoid the\nmismatch and mix pre-saved VM images without issues (slowness in\nautoscaling).\n  Thanks a lot to you all!\n  Let me know if any of you comes to Uruguay, you got free beers here!\n  Have a great day.\n\n\nEl sáb., 29 de ago. de 2020 a la(s) 01:48, Eric Wong (e@yhbt.net) escribió:\n>\n> Ivan Baldo <ibaldo@gmail.com> wrote:\n> >   Hello.\n> >   I know this is not standard usage of git, but I need a way to have\n> > more stable dates and times in the files in order to avoid rsync\n> > checksumming.\n> >   So I found this\n> > https://stackoverflow.com/questions/2179722/checking-out-old-file-with-original-create-modified-timestamps/2179876#2179876\n> > and modified it a bit to run in CentOS 7:\n> >\n> > IFS=\"\n> > \"\n> > for FILE in $(git ls-files -z | tr '\\0' '\\n')\n> > do\n> >     TIME=$(git log --pretty=format:%cd -n 1 --date=iso -- \"$FILE\")\n> >     touch -c -m -d \"$TIME\" \"$FILE\"\n> > done\n> >\n> >   Unfortunately it takes ages for a 84k files repo.\n> >   I see the CPU usage is dominated by the git log command.\n>\n> running git log for each file isn't necessary.\n>\n> On Debian, rsync actually ships the `git-set-file-times' script\n> in /usr/share/doc/rsync/scripts/ which only runs `git log' once\n> and parses it.\n>\n> You can also get my (original) version from:\n> https://yhbt.net/git-set-file-times\n>\n> >   I know a way I could use to split the work for all the CPU threads\n> > but anyway, I would like to know if you guys and girls know of a\n> > faster way to do this.\n>\n> Much of your overhead is going to be from process spawning.\n> My Perl version reduces that significantly.\n>\n> I haven't tried it with 84K files, but it'll have to keep all\n> those filenames in memory.  I'm not sure if parallelizing\n> utime() syscalls is worth it, either; maybe it helps on SSD\n> more than HDD.\n\n\n\n-- \nIvan Baldo - ibaldo@gmail.com - http://ibaldo.codigolibre.net/\nFreelance C++/PHP programmer and GNU/Linux systems administrator.\nThe sky isn't the limit!\n"}]}