git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Best way to check for a "dirty" working tree?

From
Jonathan Nieder <jrnieder@gmail.com>
Date
Jun 13, 2011, 22:22 UTC
Message-ID
<20110613222225.GA14446@elie>
In-Reply-To
<4DF381BF.3050301@dirk.my1.cc>
Hi Dirk,
Dirk Süsserott wrote:
Show 10 quoted lines
> I have a script which moves data from somewhere to my local repo and
> then checks it in, like so:
>
> -----------
> mv /tmp/foo.bar .
> git commit -am "Updated foo.bar at $timestamp"
> -----------
>
> However, before overwriting "foo.bar" in my working directory, I'd like
> to check whether my working tree is dirty (at least "foo.bar").

Interesting example. Sensible, as long as you limit the commit to foo.bar (i.e., "git commit -m ... --only foo.bar")!

Show 5 quoted lines
> I tried
>
> A) if ! git diff-index --quiet HEAD -- foo.bar; then
>        dirty=1
>    fi

To piggy-back on what Ram wrote, this is a question about the difference between porcelain (high-level) and plumbing (low-level) commands.

Generally speaking, plumbing is meant to give more stable behavior for scripts, in two ways:

 - On one hand we make a concerted effort to keep the command-line
   usage and output of plumbing stable.  By contrast, porcelain will
   change over time as we learn about the way people work.
 - On the other hand plumbing is designed to produce simple, reliable,
   and machine-friendly behavior.  For example, while "git checkout"
   will guess what the caller is trying to do based on whether its
   first argument is a branch name or a file, "git checkout-index"
   only accepts pathspecs.  Plumbing tends to produce parseable
   output and not to automatically spawn a pager when its output is
   going to the terminal or to change behavior based on configuration.

Now, a word of warning. One aspect of this "do not second-guess the caller" behavior is that low-level commands like "git diff-index" blindly trust stat() information in the index, rather than going to re-read a seemingly modified file and updating the index if the content is not changed. You can see this by running "touch foo.bar"; "git diff-index" will report the file as changed, until you use "git update-index" to refresh the stat information:

	git update-index --refresh --unmerged -q >/dev/null || :
	if ! git diff-index --quiet HEAD -- foo.bar; then
		dirty=1
	fi

Alas, this doesn't seem to be documented anywhere (except for the gitcore-tutorial(7))! It ought to be.

> Both A) and B) work. But which one is better/faster/more reliable?

I suspect the fastest (by virtue of saving a fork + exec and not having to stat files twice, once for update-index and again for diff-index) is

	git -c diff.autorefreshindex=true diff --quiet -- foo.bar

by a sad accident of history --- the "opportunistic index refresh" behavior it implements does not seem to be exposed as plumbing. If you are going to be performing such operations in a loop, then

	git update-index --refresh --unmerged -q >/dev/null || :
	for i in loop
	do
		... actions like diff-index that trust the index ...
	done

will be faster. And the latter is plumbing, with all the niceties that entails, so if I were in your shoes I'd use the latter.

Hope that helps, Jonathan

Previous: Ramkumar RamachandraNext: Dirk Süsserott
Message 3 of 4 in “Best way to check for a "dirty" working tree?”
  1. Dirk SüsserottJun 11, 2011
  2. Ramkumar RamachandraJun 12, 2011
  3. Jonathan NiederJun 13, 2011
  4. Dirk SüsserottJun 14, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.