{"thread":{"id":"27603","subject":"Best way to check for a \"dirty\" working tree?","startedAt":"2011-06-11T14:54:55Z","lastAt":"2011-06-14T13:28:07Z","messageCount":4,"participants":["Dirk Süsserott","Ramkumar Ramachandra","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"169875","messageId":"4DF381BF.3050301@dirk.my1.cc","threadId":"27603","inReplyTo":null,"subject":"Best way to check for a \"dirty\" working tree?","fromName":"Dirk Süsserott","fromEmail":"newsletter@dirk.my1.cc","sentAt":"2011-06-11T14:54:55Z","receivedAt":"2011-06-11T14:54:55Z","isPatch":false,"sender":{"key":"newsletter@dirk.my1.cc","avatar":null},"body":"Hi list,\n\nI have a script which moves data from somewhere to my local repo and\nthen checks it in, like so:\n\n-----------\nmv /tmp/foo.bar .\ngit commit -am \"Updated foo.bar at $timestamp\"\n-----------\n\nHowever, before overwriting \"foo.bar\" in my working directory, I'd like\nto check whether my working tree is dirty (at least \"foo.bar\").\n\nI tried\n\nA) if ! git diff-index --quiet HEAD -- foo.bar; then\n       dirty=1\n   fi\n\nand\n\nB) if ! git diff --quiet -- foo.bar; then\n       dirty=1\n   fi\n\nBoth A) and B) work. But which one is better/faster/more reliable? Or is\nthere a better solution? For my purpose, I cannot see a difference\nbetween diff and diff-index, except the syntax.\n\nCheers,\n    Dirk\n"},{"id":"169893","messageId":"BANLkTi=-HA1_47DvtGbVHx8twuEAxT8STQ@mail.gmail.com","threadId":"27603","inReplyTo":"4DF381BF.3050301@dirk.my1.cc","subject":"Re: Best way to check for a \"dirty\" working tree?","fromName":"Ramkumar Ramachandra","fromEmail":"artagnon@gmail.com","sentAt":"2011-06-12T12:23:13Z","receivedAt":"2011-06-12T12:23:13Z","isPatch":false,"sender":{"key":"r@artagnon.com","avatar":"https://avatars.githubusercontent.com/u/37226?v=4"},"body":"Hi Dirk,\n\nDirk Süsserott writes:\n> A) if ! git diff-index --quiet HEAD -- foo.bar; then\n>       dirty=1\n>   fi\n>\n> and\n>\n> B) if ! git diff --quiet -- foo.bar; then\n>       dirty=1\n>   fi\n>\n> Both A) and B) work. But which one is better/faster/more reliable? Or is\n> there a better solution? For my purpose, I cannot see a difference\n> between diff and diff-index, except the syntax.\n\ndiff is a more porcelain'ish command, while diff-index is closer to\nthe plumbing.  Therefore, diff contains some extra argument parsing/\npretty printing code that your script doesn't utilize -- use\ndiff-index.  Also, look at the various scripts in git.git to see what\nthey use; for example, require_clean_work_tree in git-sh-setup.sh.\n\n-- Ram\n"},{"id":"169968","messageId":"20110613222225.GA14446@elie","threadId":"27603","inReplyTo":"4DF381BF.3050301@dirk.my1.cc","subject":"Re: Best way to check for a \"dirty\" working tree?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2011-06-13T22:22:48Z","receivedAt":"2011-06-13T22:22:48Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi Dirk,\n\nDirk Süsserott wrote:\n\n> I have a script which moves data from somewhere to my local repo and\n> then checks it in, like so:\n>\n> -----------\n> mv /tmp/foo.bar .\n> git commit -am \"Updated foo.bar at $timestamp\"\n> -----------\n>\n> However, before overwriting \"foo.bar\" in my working directory, I'd like\n> to check whether my working tree is dirty (at least \"foo.bar\").\n\nInteresting example.  Sensible, as long as you limit the commit to\nfoo.bar (i.e., \"git commit -m ... --only foo.bar\")!\n\n> I tried\n>\n> A) if ! git diff-index --quiet HEAD -- foo.bar; then\n>        dirty=1\n>    fi\n\nTo piggy-back on what Ram wrote, this is a question about the\ndifference between porcelain (high-level) and plumbing (low-level)\ncommands.\n\nGenerally speaking, plumbing is meant to give more stable behavior for\nscripts, in two ways:\n\n - On one hand we make a concerted effort to keep the command-line\n   usage and output of plumbing stable.  By contrast, porcelain will\n   change over time as we learn about the way people work.\n\n - On the other hand plumbing is designed to produce simple, reliable,\n   and machine-friendly behavior.  For example, while \"git checkout\"\n   will guess what the caller is trying to do based on whether its\n   first argument is a branch name or a file, \"git checkout-index\"\n   only accepts pathspecs.  Plumbing tends to produce parseable\n   output and not to automatically spawn a pager when its output is\n   going to the terminal or to change behavior based on configuration.\n\nNow, a word of warning.  One aspect of this \"do not second-guess the\ncaller\" behavior is that low-level commands like \"git diff-index\"\nblindly trust stat() information in the index, rather than going to\nre-read a seemingly modified file and updating the index if the\ncontent is not changed.  You can see this by running \"touch foo.bar\";\n\"git diff-index\" will report the file as changed, until you use \"git\nupdate-index\" to refresh the stat information:\n\n\tgit update-index --refresh --unmerged -q >/dev/null || :\n\tif ! git diff-index --quiet HEAD -- foo.bar; then\n\t\tdirty=1\n\tfi\n\nAlas, this doesn't seem to be documented anywhere (except for the\ngitcore-tutorial(7))!  It ought to be.\n\n> Both A) and B) work. But which one is better/faster/more reliable?\n\nI suspect the fastest (by virtue of saving a fork + exec and not\nhaving to stat files twice, once for update-index and again for\ndiff-index) is\n\n\tgit -c diff.autorefreshindex=true diff --quiet -- foo.bar\n\nby a sad accident of history --- the \"opportunistic index refresh\"\nbehavior it implements does not seem to be exposed as plumbing.\nIf you are going to be performing such operations in a loop, then\n\n\tgit update-index --refresh --unmerged -q >/dev/null || :\n\tfor i in loop\n\tdo\n\t\t... actions like diff-index that trust the index ...\n\tdone\n\nwill be faster.  And the latter is plumbing, with all the niceties\nthat entails, so if I were in your shoes I'd use the latter.\n\nHope that helps,\nJonathan\n"},{"id":"170002","messageId":"4DF761E7.8040707@dirk.my1.cc","threadId":"27603","inReplyTo":"20110613222225.GA14446@elie","subject":"Re: Best way to check for a \"dirty\" working tree?","fromName":"Dirk Süsserott","fromEmail":"newsletter@dirk.my1.cc","sentAt":"2011-06-14T13:28:07Z","receivedAt":"2011-06-14T13:28:07Z","isPatch":false,"sender":{"key":"newsletter@dirk.my1.cc","avatar":null},"body":"Hi Jonathan,\n\nAm 14.06.2011 00:22 schrieb Jonathan Nieder:\n> Hi Dirk,\n> \n> Dirk Süsserott wrote:\n> \n>> I have a script which moves data from somewhere to my local repo and\n>> then checks it in, like so:\n>>\n>> -----------\n>> mv /tmp/foo.bar .\n>> git commit -am \"Updated foo.bar at $timestamp\"\n>> -----------\n>>\n>> However, before overwriting \"foo.bar\" in my working directory, I'd like\n>> to check whether my working tree is dirty (at least \"foo.bar\").\n> \n> Interesting example.  Sensible, as long as you limit the commit to\n> foo.bar (i.e., \"git commit -m ... --only foo.bar\")!\n\nUhh, nice hint. I didn't know that git-commit accepts a path, too.\nThat's safer. However, in my particular case the working tree is either\nclean or exactly the file in question has changed. If sth. else changes\n(e.g. my commit-script) I do that in a separate \"transaction\".\n\n> Now, a word of warning.  One aspect of this \"do not second-guess the\n> caller\" behavior is that low-level commands like \"git diff-index\"\n> blindly trust stat() information in the index, rather than going to\n> re-read a seemingly modified file and updating the index if the\n> content is not changed.  You can see this by running \"touch foo.bar\";\n> \"git diff-index\" will report the file as changed, until you use \"git\n> update-index\" to refresh the stat information:\n> \n> \tgit update-index --refresh --unmerged -q >/dev/null || :\n> \tif ! git diff-index --quiet HEAD -- foo.bar; then\n> \t\tdirty=1\n> \tfi\n> \n> Alas, this doesn't seem to be documented anywhere (except for the\n> gitcore-tutorial(7))!  It ought to be.\n\nHmm, it MUST be documented somewhere, because I have several scripts\nthat use \"update-index --refresh\" to get rid of what I call \"phantom\nchanges\": sometimes I transfer (scp) files from a remote machine to the\nlocal tree. The set of files is already known to Git, so my first guess\nwas that Gitk would only show the \"real\" diff, but it actually showed\n*all* transferred files as changed. After running \"git status\" Gitk does\nit right and shows only content's diff. Surprisingly, \"git status\" seems\nto be a read/write operation and does \"update-index --refresh\" in the\nbackground. After some research I learned about \"update-index --refresh\"\nand use it frequently for scp'ed files.\n\nUnfortunately, I cannot remember *where* I learned about it.\n\n> Hope that helps,\n> Jonathan\n\nThat helped a lot. Thank you,\nDirk\n"}]}