git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: 'git status' is not read-only fs friendly

From
Junio C Hamano <junkio@cox.net>
Date
Feb 11, 2007, 06:33 UTC
Message-ID
<7vtzxtdwz9.fsf@assigned-by-dhcp.cox.net>
In-Reply-To
<Pine.LNX.4.64.0702100913020.8424@woody.linux-foundation.org>
Linus Torvalds <torvalds@linux-foundation.org> writes:
Show 22 quoted lines
> On Sat, 10 Feb 2007, Nicolas Pitre wrote:
>> > >
>> > > Because git-status itself is conceptually a read-only operation, and 
>> > > having it barf on a read-only file system is justifiably a bug.
>> > 
>> > I do not 100% agree that it is conceptually a read-only operation.
>> 
>> It is.
>
> It really isn't. 
>
> It's not even a "technical issue". It's a fundamental optimization. Sure, 
> you can call optimizations just "technical issues", but the fact is, it's
> one of the things that makes git so _usable_ on large archives. At some 
> point, an "optimization" is no longer just about making things slightly 
> faster, it's about something much bigger, and has real semantic meaning.
> ...
> THIS IS NOT "JUST A TECHNICAL ISSUE". 
> ...
> And the index is what makes it so. 
>
> And that's why it's important to keep the index up-to-date.
I think a one paragraph summary of your argument is:
 - index is a good thing -- it is what makes the difference
   between usable and unusable.
 - git-status needs to refresh the index in order to do its
   thing efficiently and usably _anyway_, so once it spends
   cycles to do so, it is senseless not to write the refreshed
   index out when it can.

I do not think anybody disputes that in a repository with 20k+ paths, it is sensible to leave the index stat-dirty for all paths. But I think your example

	read-tree HEAD

misses the point by stressing the importance of index too much. Index is important for the usability and I do not think anybody is disputing it.

The thing is, nobody switches the index that way without running "update-index --refresh" afterwards. Normal people would use git-reset to switch to a different tree object, and the command does that for you. If you are a hardcore, you would know to use "read-tree -m HEAD" at least to avoid making paths unnecessarily stat-dirty. Your example, while it is valid and demonstrates why the index is a good thing very well, is simply not part of a normal workflow and not very relevant when discussing the performance ramifications of what state "git-status" should leave the index in.

When I said "calling 'update-index --refresh' in git-status loses stat-dirtiness information", I was certainly _NOT_ talking about losing the information that 20k+ paths used to be stat-dirty because the user did "read-tree HEAD" earlier.

At least for me, it is very normal to do something like this.
 * start from a clean index.
 * edit cache.h, diff.h, and diff-lib.c.
 * stop, think, and realize that my earlier edit to change one
   function prototype in diff.h was not needed, and revert the
   change to that line still in the editor.
 * fix things up further by editing other files.

And then, I would run "git diff" to see where I am. I still remember that I touched diff.h and I also remember that I once changed a function prototype but then decided the change was not necessary after all, but I do not remember if I changed anything else in the file. It is _very_ assuring to see the emptiness that follows "git diff --git" header for diff.h in such a case. Seeing the path to be stat-dirty is a very good thing for me, because otherwise I might lose a few seconds thinking that what I thought I touched might have been cache.h and not diff.h.

To me, running "git status" is "wrapping things up" step. I do not need that stat-dirty assurance "git diff" gave me at that point. Not seeing diff.h in "modified but updated" list is a good thing. And in my workflow, after that 'wrapping things up" step, I do not need that stat-dirty assurance _anymore_.

I think Nico is correct to point out that "not _anymore_" part of the above reasoning of mine assumes _my_ workflow and preference, and I think that is a valid point. Not saving the refreshed index would make the stat-dirtiness for diff.h to come back, which would be inconvenient and annoying to me.

But the user might want to keep it stat-dirty after running "git-status". People in "not _anymore_" camp like me can throw the stat-dirtiness away by "update-index --refresh". I do not think he (or anybody) is advocating to keep 20k+ paths in stat-dirty state (arguably, "artificially" due to use of "read-tree HEAD"), so your example using "read-tree HEAD" only confuses the discussion.

Having said all that, I do agree with you that git-status should throw that stat-dirtiness information away by saving the refreshed index. Doing otherwise is annoying to me as I already said, and I do not think of a valid reason for the user to want to keep stat-dirtiness information after running "git-status", because to me the whole point of running "git-status" is to start wrapping things up.

Previous: Nicolas PitreNext: Shawn O. Pearce
Message 48 of 52 in “'git status' is not read-only fs friendly”
  1. Marco CostalbaFeb 9, 2007
  2. Linus TorvaldsFeb 9, 2007
  3. Marco CostalbaFeb 9, 2007
  4. Junio C HamanoFeb 9, 2007
  5. Junio C HamanoFeb 9, 2007
  6. Morten WelinderFeb 9, 2007
  7. Theodore TsoFeb 9, 2007
  8. Marco CostalbaFeb 9, 2007
  9. Linus TorvaldsFeb 9, 2007
  10. Junio C HamanoFeb 10, 2007
  11. Junio C HamanoFeb 10, 2007
  12. 1/2 run_diff_{files,index}(): update calling convention.Junio C Hamano, Feb 10, 2007
  13. Marco CostalbaFeb 10, 2007
  14. Junio C HamanoFeb 10, 2007
  15. Marco CostalbaFeb 10, 2007
  16. Marco CostalbaFeb 10, 2007
  17. Junio C HamanoFeb 10, 2007
  18. Marco CostalbaFeb 10, 2007
  19. Junio C HamanoFeb 10, 2007
  20. Marco CostalbaFeb 10, 2007
  21. 2/2 git-runstatus --refreshJunio C Hamano, Feb 10, 2007
  22. Johannes SchindelinFeb 10, 2007
  23. Marco CostalbaFeb 10, 2007
  24. Johannes SchindelinFeb 10, 2007
  25. Marco CostalbaFeb 10, 2007
  26. Marco CostalbaFeb 10, 2007
  27. Junio C HamanoFeb 10, 2007
  28. Johannes SchindelinFeb 10, 2007
  29. Junio C HamanoFeb 11, 2007
  30. Johannes SchindelinFeb 11, 2007
  31. Johannes SchindelinFeb 11, 2007
  32. Junio C HamanoFeb 11, 2007
  33. Johannes SchindelinFeb 11, 2007
  34. Johannes SchindelinFeb 10, 2007
  35. Marco CostalbaFeb 10, 2007
  36. Nicolas PitreFeb 10, 2007
  37. Junio C HamanoFeb 10, 2007
  38. Nicolas PitreFeb 10, 2007
  39. Junio C HamanoFeb 10, 2007
  40. Nicolas PitreFeb 10, 2007
  41. Junio C HamanoFeb 10, 2007
  42. Theodore TsoFeb 10, 2007
  43. Nicolas PitreFeb 10, 2007
  44. Theodore TsoFeb 10, 2007
  45. Marco CostalbaFeb 10, 2007
  46. Linus TorvaldsFeb 10, 2007
  47. Nicolas PitreFeb 10, 2007
  48. Junio C HamanoFeb 11, 2007
  49. Shawn O. PearceFeb 11, 2007
  50. Johannes SchindelinFeb 10, 2007
  51. Junio C HamanoFeb 10, 2007
  52. Marco CostalbaFeb 10, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.