{"thread":{"id":"16417","subject":"Bad git status performance","startedAt":"2008-11-21T00:28:14Z","lastAt":"2008-11-21T20:07:16Z","messageCount":5,"participants":["Jean-Luc Herren","David Bryson","Michael J Gruber"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"96295","messageId":"4926009E.4040203@gmx.ch","threadId":"16417","inReplyTo":null,"subject":"Bad git status performance","fromName":"Jean-Luc Herren","fromEmail":"jlh@gmx.ch","sentAt":"2008-11-21T00:28:14Z","receivedAt":"2008-11-21T00:28:14Z","isPatch":false,"sender":{"key":"jlh@gmx.ch","avatar":null},"body":"Hi list!\n\nI'm getting bad performance on 'git status' when I have staged\nmany changes to big files.  For example, consider this:\n\n$ git init\nInitialized empty Git repository in $HOME/test/.git/\n\n$ for X in $(seq 100); do dd if=/dev/zero of=$X bs=1M count=1 2> /dev/null; done\n\n$ git add .\n\n$ git commit -m 'Lots of zeroes'\nCreated initial commit ed54346: Lots of zeroes\n 100 files changed, 0 insertions(+), 0 deletions(-)\n create mode 100644 1\n create mode 100644 10\n...\n create mode 100644 98\n create mode 100644 99\n\n$ for X in $(seq 100); do echo > $X; done\n\n$ time git status\n# On branch master\n# Changed but not updated:\n#   (use \"git add <file>...\" to update what will be committed)\n#\n#       modified:   1\n#       modified:   10\n...\n#       modified:   98\n#       modified:   99\n#\nno changes added to commit (use \"git add\" and/or \"git commit -a\")\n\nreal    0m0.003s\nuser    0m0.001s\nsys     0m0.002s\n\n$ git add -u\n\n$ time git status\n# On branch master\n# Changes to be committed:\n#   (use \"git reset HEAD <file>...\" to unstage)\n#\n#       modified:   1\n#       modified:   10\n...\n#       modified:   98\n#       modified:   99\n#\n\nreal    0m16.291s\nuser    0m16.054s\nsys     0m0.221s\n\nThe first 'git status' shows the same difference as the second,\njust the second time it's staged instead of unstaged.  Why does it\ntake 16 seconds the second time when it's instant the first time?\n\n(Side note: There once was a discussion about adding natural order\nof branch names, but seems it never made it into git.  The same\nwould make sense for 'git status' too.)\n\nCheers,\njlh\n"},{"id":"96297","messageId":"20081121004242.GD6458@eratosthenes.cryptobackpack.org","threadId":"16417","inReplyTo":"4926009E.4040203@gmx.ch","subject":"Re: Bad git status performance","fromName":"David Bryson","fromEmail":"david@statichacks.org","sentAt":"2008-11-21T00:42:42Z","receivedAt":"2008-11-21T00:42:42Z","isPatch":false,"sender":{"key":"david@statichacks.org","avatar":"https://gravatar.com/avatar/b8796a0b286799d99dcbaea3fd3e8675648cc094ff10b07bae8fe3bc0ac40b9c?d=mp&s=160"},"body":"Hi,\n\nOn Fri, Nov 21, 2008 at 01:28:14AM +0100 or thereabouts, Jean-Luc Herren wrote:\n> Hi list!\n> \n> I'm getting bad performance on 'git status' when I have staged\n> many changes to big files.  For example, consider this:\n> \n[snip]\n> $ time git status\n> # On branch master\n> # Changes to be committed:\n> #   (use \"git reset HEAD <file>...\" to unstage)\n> #\n> #       modified:   1\n> #       modified:   10\n> ...\n> #       modified:   98\n> #       modified:   99\n> #\n> \n> real    0m16.291s\n> user    0m16.054s\n> sys     0m0.221s\n> \n> The first 'git status' shows the same difference as the second,\n> just the second time it's staged instead of unstaged.  Why does it\n> take 16 seconds the second time when it's instant the first time?\n\nI had similar problems with a repository that contained several tarballs\nof gcc and the linux kernel(don't ask me why it was not my repository).\n\nSome weeks ago I mentioned this on IRC, and the problem really was not\nnecessarily git.  The way it was explained to me(and please correct or\nclairify where I am wrong) is that git asked linux for the status of\nthose files and being that they are so large they were swapped out of\nmemory.\n\nThe result is the kernel reading those large files back in to see if\nthey have changed at all.  My impression is that this is not a git bug\nbut a cache-tuning problem.\n\nDave\n"},{"id":"96334","messageId":"4926ADB8.5000307@gmx.ch","threadId":"16417","inReplyTo":"c9e534200811201711y887ddd2t33013ec4a7db3c9a@mail.gmail.com","subject":"Re: Bad git status performance","fromName":"Jean-Luc Herren","fromEmail":"jlh@gmx.ch","sentAt":"2008-11-21T12:46:48Z","receivedAt":"2008-11-21T12:46:48Z","isPatch":false,"sender":{"key":"jlh@gmx.ch","avatar":null},"body":"Glenn Griffin wrote:\n> On Thu, Nov 20, 2008 at 4:28 PM, Jean-Luc Herren <jlh@gmx.ch> wrote:\n>> The first 'git status' shows the same difference as the second,\n>> just the second time it's staged instead of unstaged.  Why does it\n>> take 16 seconds the second time when it's instant the first time?\n> \n> I believe the two runs of git status need to do very different things.\n>  When run the first time, git knows the files in your working\n> directory are not in the index so it can easily say those files are\n> 'Changed but not updated' just from their existence.\n\nI might be mistaken about how the index works, but those paths\n*are* in the index at that time.  They just have the old content,\ni.e. the same content as in HEAD.  When HEAD == index, then\nnothing is staged.\n\nBut the presence of those files alone doesn't tell you that they\nhave changed.  You have to look at the content and compare it to\nthe index (== HEAD in this situation) to see whether they have\nchanged or not and for some reason git can do this very quickly.\n\n> The second run\n> those files do exist in both the index and the working directory, so\n> git status first shows the files that are 'Changes to be committed'\n> and that should be fast, but additionally git status will check to see\n> if those files in your working directory have changed since you added\n> them to the index.\n\nWhich is basically the same comparision as above, just it turns\nout that they have not changed.  But even then, we're talking\nabout comparing a 1 byte file in the index to a 1 byte file in the\nwork tree.  That doesn't take 16 seconds, even for 100 files.\n\nSo this makes me believe it's the first step (comparing HEAD to\nthe index to show staged changes) that is slow.  And when you\ncompare a 1MB file to a 1 byte file, you don't need to read all of\nthe big file, you can tell they're not the same right after the\nfirst byte.  (Even an doing stat() is enough, since the size is\nnot the same.)\n\nAnother thing that came to my mind is maybe rename detection kicks\nin, even though no path vanished and none is new.  I believe\nrename detection doesn't happen for unstaged changes, which might\nexplain the difference in speed.\n\nbtw, I forgot to mention that I get this with branches maint,\nmaster, next and pu.\n\n(And I hope you don't mind I take this back to the list.)\n\njlh\n"},{"id":"96345","messageId":"4926D196.3000301@drmicha.warpmail.net","threadId":"16417","inReplyTo":"4926ADB8.5000307@gmx.ch","subject":"Re: Bad git status performance","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2008-11-21T15:19:50Z","receivedAt":"2008-11-21T15:19:50Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Jean-Luc Herren venit, vidit, dixit 21.11.2008 13:46:\n> Glenn Griffin wrote:\n>> On Thu, Nov 20, 2008 at 4:28 PM, Jean-Luc Herren <jlh@gmx.ch> wrote:\n>>> The first 'git status' shows the same difference as the second,\n>>> just the second time it's staged instead of unstaged.  Why does it\n>>> take 16 seconds the second time when it's instant the first time?\n>> I believe the two runs of git status need to do very different things.\n>>  When run the first time, git knows the files in your working\n>> directory are not in the index so it can easily say those files are\n>> 'Changed but not updated' just from their existence.\n> \n> I might be mistaken about how the index works, but those paths\n> *are* in the index at that time.  They just have the old content,\n> i.e. the same content as in HEAD.  When HEAD == index, then\n> nothing is staged.\n> \n> But the presence of those files alone doesn't tell you that they\n> have changed.  You have to look at the content and compare it to\n> the index (== HEAD in this situation) to see whether they have\n> changed or not and for some reason git can do this very quickly.\n> \n>> The second run\n>> those files do exist in both the index and the working directory, so\n>> git status first shows the files that are 'Changes to be committed'\n>> and that should be fast, but additionally git status will check to see\n>> if those files in your working directory have changed since you added\n>> them to the index.\n> \n> Which is basically the same comparision as above, just it turns\n> out that they have not changed.  But even then, we're talking\n> about comparing a 1 byte file in the index to a 1 byte file in the\n> work tree.  That doesn't take 16 seconds, even for 100 files.\n> \n> So this makes me believe it's the first step (comparing HEAD to\n> the index to show staged changes) that is slow.  And when you\n> compare a 1MB file to a 1 byte file, you don't need to read all of\n> the big file, you can tell they're not the same right after the\n> first byte.  (Even an doing stat() is enough, since the size is\n> not the same.)\n> \n> Another thing that came to my mind is maybe rename detection kicks\n> in, even though no path vanished and none is new.  I believe\n> rename detection doesn't happen for unstaged changes, which might\n> explain the difference in speed.\n> \n> btw, I forgot to mention that I get this with branches maint,\n> master, next and pu.\n\nInterestingly, all of\n\ngit diff --stat\ngit diff --stat --cached\ngit diff --stat HEAD\n\nare \"fast\" (0.2s or so), i.e. diffing index-wtree, HEAD-index,\nHEAD-wtree. Linus' threaded stat doesn't help either for status, btw (20s).\n\nExperimenting further: Using 10 files with 10MB each (rather than 100\ntimes 1MB) brings down the time by a factor 10 roughly - and so does\nusing 100 files with 100k each. Huh? Latter may be expected (10MB\ntotal), but former (100MB total)?\n\nNow it's getting funny: Changing your \"echo >\" to \"echo \">>\" (in your\n100 files 1MB case) makes things \"almost fast\" again (1.3s).\n\nOK, it's \"use the source, Luke\" time... Actually the part you don't see\ntakes the most time:\nwt_status_print_updated()\n\nAnd in fact I can confirm your suspicion: wt_status_print_updated()\nenforces rename detection (ignoring any config). Forcing it off\n(rev.diffopt.detect_rename = 0;) cuts down the 20s to 0.75s.\n\nHow about a config option status.renames (or something like -M) for status?\n\nMichael\n"},{"id":"96351","messageId":"492714F4.1090807@gmx.ch","threadId":"16417","inReplyTo":"4926D196.3000301@drmicha.warpmail.net","subject":"Re: Bad git status performance","fromName":"Jean-Luc Herren","fromEmail":"jlh@gmx.ch","sentAt":"2008-11-21T20:07:16Z","receivedAt":"2008-11-21T20:07:16Z","isPatch":false,"sender":{"key":"jlh@gmx.ch","avatar":null},"body":"Michael J Gruber wrote:\n> Experimenting further: Using 10 files with 10MB each (rather than 100\n> times 1MB) brings down the time by a factor 10 roughly - and so does\n> using 100 files with 100k each. Huh? Latter may be expected (10MB\n> total), but former (100MB total)?\n\n100 files at each 100k gives me 1.73s, so about 10x speed up.  So\nit seems git indeed looks at the content of the files and having a\ntenth of the content means it's ten times as fast.\n\nInterestingly, using only a single file of 100MB gives me 0.6s.\nWhich is still very slow for the job of telling that a 100MB file\nis not equal to a 1 byte file.  And certainly there's no renaming\ngoing on with a single file.\n\n> Now it's getting funny: Changing your \"echo >\" to \"echo \">>\" (in your\n> 100 files 1MB case) makes things \"almost fast\" again (1.3s).\n\nSame here and that's pretty interesting, because in this situation\nI can understand the slow down: Comparing two 1MB files that\ndiffer only at their ends is expected to take some time, as you\nhave to go through the entire file until you notice they're not\nthe same.\n\njlh\n"}]}