{"thread":{"id":"16014","subject":"git performance","startedAt":"2008-10-22T20:17:16Z","lastAt":"2008-10-24T23:10:09Z","messageCount":22,"participants":["Edward Ned Harvey","Jeff King","Peter Harris","Jakub Narebski","Andreas Ericsson","Matthieu Moy","Nguyen Thai Ngoc Duy","Daniel Barkalow","Nanako Shiraishi","Pete Harlan","George Shammas","Linus Torvalds"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"93720","messageId":"000801c93483$2fdad340$8f9079c0$@com","threadId":"16014","inReplyTo":null,"subject":"git performance","fromName":"Edward Ned Harvey","fromEmail":"git@nedharvey.com","sentAt":"2008-10-22T20:17:16Z","receivedAt":"2008-10-22T20:17:16Z","isPatch":false,"sender":{"key":"git@nedharvey.com","avatar":null},"body":"I see things all over the Internet saying git is fast.  I'm currently struggling with poor svn performance and poor attitude of svn developers, so I'd like to consider switching to git.  A quick question first.\n\nThe core of the performance problem I'm facing is the need to \"walk the tree\" for many thousand files.  Every time I do \"svn update\" or \"svn status\" the svn client must stat every file to check for local modifications (a coffee cup or a beer worth of stats).  In essence, this is unavoidable if there is no mechanism to constantly monitor filesystem activity during normal operations.  Analogous to filesystem journaling.\n\nSo - I didn't see anything out there saying \"git is fast because it uses inotify\" or anything like that.  Perhaps git would not help me at all?  Because git still needs to stat all the files in the tree?\n"},{"id":"93723","messageId":"20081022203624.GA4585@coredump.intra.peff.net","threadId":"16014","inReplyTo":"000801c93483$2fdad340$8f9079c0$@com","subject":"Re: git performance","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-10-22T20:36:25Z","receivedAt":"2008-10-22T20:36:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Oct 22, 2008 at 04:17:16PM -0400, Edward Ned Harvey wrote:\n\n> So - I didn't see anything out there saying \"git is fast because it\n> uses inotify\" or anything like that.  Perhaps git would not help me at\n> all?  Because git still needs to stat all the files in the tree?\n\nYes, it does stat all the files. How many files are you talking about,\nand what platform?  From a warm cache on Linux, the 23,000 files kernel\nrepo takes about a tenth of a second to stat all files for me (and this\non a several year-old machine). And of course many operations don't\nrequire stat'ing at all (like looking at logs, or diffs that don't\ninvolve the working tree).\n\n-Peff\n"},{"id":"93730","messageId":"eaa105840810221413m4d0ed51ejab28f66493a12a13@mail.gmail.com","threadId":"16014","inReplyTo":"20081022203624.GA4585@coredump.intra.peff.net","subject":"Re: git performance","fromName":"Peter Harris","fromEmail":"git@peter.is-a-geek.org","sentAt":"2008-10-22T21:13:56Z","receivedAt":"2008-10-22T21:13:56Z","isPatch":false,"sender":{"key":"git@peter.is-a-geek.org","avatar":null},"body":"On Wed, Oct 22, 2008 at 4:36 PM, Jeff King wrote:\n> On Wed, Oct 22, 2008 at 04:17:16PM -0400, Edward Ned Harvey wrote:\n>\n>> So - I didn't see anything out there saying \"git is fast because it\n>> uses inotify\" or anything like that.  Perhaps git would not help me at\n>> all?  Because git still needs to stat all the files in the tree?\n>\n> Yes, it does stat all the files. How many files are you talking about,\n> and what platform?  From a warm cache on Linux, the 23,000 files kernel\n> repo takes about a tenth of a second to stat all files for me (and this\n> on a several year-old machine). And of course many operations don't\n> require stat'ing at all (like looking at logs, or diffs that don't\n> involve the working tree).\n\nWindows is rather slower than Linux, so differences are more obvious.\nI find git feels \"only\" about 2x as fast as svn at status. svn has to\nstat all of its base files too, whereas git has the index. git pull\n(vs svn update) feels better than 2x faster, since git doesn't need to\nwalk the tree and lock every sub-dir before it even connects to the\nremote server.\n\nSo we're not talking 'inotify' fast, but maybe half a cup of coffee\ninstead of a full cup if you have that many files.\n\n\"git-svn\" is really quite good. I recommend you try a quick (trunk and\nmaybe one branch only, last few revisions only) import of your svn\ntree to test with.\n\nPeter Harris\n"},{"id":"93734","messageId":"000901c93490$e0c40ed0$a24c2c70$@com","threadId":"16014","inReplyTo":"20081022203624.GA4585@coredump.intra.peff.net","subject":"RE: git performance","fromName":"Edward Ned Harvey","fromEmail":"git@nedharvey.com","sentAt":"2008-10-22T21:55:14Z","receivedAt":"2008-10-22T21:55:14Z","isPatch":false,"sender":{"key":"git@nedharvey.com","avatar":null},"body":"> Yes, it does stat all the files. How many files are you talking about,\n> and what platform?  From a warm cache on Linux, the 23,000 files kernel\n> repo takes about a tenth of a second to stat all files for me (and this\n> on a several year-old machine). And of course many operations don't\n> require stat'ing at all (like looking at logs, or diffs that don't\n> involve the working tree).\n\nNo worries.  No solution can meet everyone's needs.\n\nI'm talking about 40-50,000 files, on multi-user production linux, which means the cache is never warm, except when I'm benchmarking.  Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn st\" takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me, performance was approx the same for files on local disk versus NFS.  Probably the best solution for us is perforce, we just don't like the pricetag.\n\nOut of curiosity, what are they talking about, when they say \"git is fast?\"  Just the fact that it's all local disk, or is there more to it than that?  I could see - git would probably outperform perforce for versioning of large files (let's say iso files) to benefit from sustained local disk IO, while perforce would probably outperform anything I can think of, operating on thousands of tiny files, because it will never walk the tree.\n"},{"id":"93741","messageId":"m3d4hsi708.fsf@localhost.localdomain","threadId":"16014","inReplyTo":"000801c93483$2fdad340$8f9079c0$@com","subject":"Re: git performance","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-10-22T22:42:51Z","receivedAt":"2008-10-22T22:42:51Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Edward Ned Harvey\" <git@nedharvey.com> writes:\n\n> I see things all over the Internet saying git is fast.  I'm\n> currently struggling with poor svn performance and poor attitude of\n> svn developers, so I'd like to consider switching to git.  A quick\n> question first.\n> \n> The core of the performance problem I'm facing is the need to \"walk\n> the tree\" for many thousand files.  Every time I do \"svn update\" or\n> \"svn status\" the svn client must stat every file to check for local\n> modifications (a coffee cup or a beer worth of stats).  In essence,\n> this is unavoidable if there is no mechanism to constantly monitor\n> filesystem activity during normal operations.  Analogous to\n> filesystem journaling.\n> \n> So - I didn't see anything out there saying \"git is fast because it\n> uses inotify\" or anything like that.  Perhaps git would not help me\n> at all?  Because git still needs to stat all the files in the tree?\n\nhttp://git.or.cz/gitwiki/GitBenchmarks\n\nWhile it should be possible to use 'assume unchanged' bit together\nwith inotify / icron, it is not something tha is done; IIRC Mercurial\nhad Linux-only InotifyPlugin...\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"93766","messageId":"49002399.20303@op5.se","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"Re: git performance","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-10-23T07:11:21Z","receivedAt":"2008-10-23T07:11:21Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Edward Ned Harvey wrote:\n>> Yes, it does stat all the files. How many files are you talking about,\n>> and what platform?  From a warm cache on Linux, the 23,000 files kernel\n>> repo takes about a tenth of a second to stat all files for me (and this\n>> on a several year-old machine). And of course many operations don't\n>> require stat'ing at all (like looking at logs, or diffs that don't\n>> involve the working tree).\n> \n> No worries.  No solution can meet everyone's needs.\n> \n> I'm talking about 40-50,000 files, on multi-user production linux, which means the cache is never warm, except when I'm benchmarking.  Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn st\" takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me, performance was approx the same for files on local disk versus NFS.  Probably the best solution for us is perforce, we just don't like the pricetag.\n> \n> Out of curiosity, what are they talking about, when they say \"git is fast?\"  Just the fact that it's all local disk, or is there more to it than that?  I could see - git would probably outperform perforce for versioning of large files (let's say iso files) to benefit from sustained local disk IO, while perforce would probably outperform anything I can think of, operating on thousands of tiny files, because it will never walk the tree.\n> \n\n\n\n> \n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"93767","messageId":"490023A0.4080901@op5.se","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"Re: git performance","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-10-23T07:11:28Z","receivedAt":"2008-10-23T07:11:28Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Edward Ned Harvey wrote:\n>> Yes, it does stat all the files. How many files are you talking about,\n>> and what platform?  From a warm cache on Linux, the 23,000 files kernel\n>> repo takes about a tenth of a second to stat all files for me (and this\n>> on a several year-old machine). And of course many operations don't\n>> require stat'ing at all (like looking at logs, or diffs that don't\n>> involve the working tree).\n> \n> No worries.  No solution can meet everyone's needs.\n> \n> I'm talking about 40-50,000 files, on multi-user production linux, which means the cache is never warm, except when I'm benchmarking.  Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn st\" takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me, performance was approx the same for files on local disk versus NFS.  Probably the best solution for us is perforce, we just don't like the pricetag.\n> \n> Out of curiosity, what are they talking about, when they say \"git is fast?\"  Just the fact that it's all local disk, or is there more to it than that?  I could see - git would probably outperform perforce for versioning of large files (let's say iso files) to benefit from sustained local disk IO, while perforce would probably outperform anything I can think of, operating on thousands of tiny files, because it will never walk the tree.\n> \n> \n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"93761","messageId":"49002AA1.80203@op5.se","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"Re: git performance","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-10-23T07:41:21Z","receivedAt":"2008-10-23T07:41:21Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Edward Ned Harvey wrote:\n>> Yes, it does stat all the files. How many files are you talking\n>> about, and what platform?  From a warm cache on Linux, the 23,000\n>> files kernel repo takes about a tenth of a second to stat all files\n>> for me (and this on a several year-old machine). And of course many\n>> operations don't require stat'ing at all (like looking at logs, or\n>> diffs that don't involve the working tree).\n> \n> No worries.  No solution can meet everyone's needs.\n> \n> I'm talking about 40-50,000 files, on multi-user production linux,\n\nUmm... using git to track a production server? I think there's something\nin your specific use-case that eluded pretty much everyone here the\nfirst time you asked about it.\n\ngit was built to maintain the linux kernel with its patch-and-merge based\nworkflow, 117k commits and 25k files. It's *good* at that sort of thing,\nbut a lot of features are \"source-code management\" specific. It sounds to\nme you're asking for something that will keep a backup of most of your\nentire system (apart from /home), which it's not really suited for. For\ninstance, it doesn't keep track of mode-bits on files (apart from\n\"executable or not\").\n\n> which means the cache is never warm, except when I'm benchmarking.\n> Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn st\"\n> takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me,\n> performance was approx the same for files on local disk versus NFS.\n> Probably the best solution for us is perforce, we just don't like the\n> pricetag.\n> \n> Out of curiosity, what are they talking about, when they say \"git is\n> fast?\"\n\nMerges, patch application, committing, history walking and data\ntransfers are all extremely quick operations under git.\n\nActually, history walking isn't extremely quick, but several neat\ntricks are in place that make it *seem* quick. Running\n\"git log drivers/net/wireless\" on the linux kernel with a cold\ncache starts spitting out output after about 1 second on my measly\nlaptop (where the kernel has 117k commits on 25k files).\n\n>  Just the fact that it's all local disk, or is there more to\n> it than that?  I could see - git would probably outperform perforce\n> for versioning of large files (let's say iso files) to benefit from\n> sustained local disk IO, while perforce would probably outperform\n> anything I can think of, operating on thousands of tiny files,\n> because it will never walk the tree.\n> \n\nGit doesn't *have* to walk the tree either. \"git status\" obviously\nhas to do that, since you're asking \"what files have changed in this\ntree since I last added stuff to the index\", but you can use git just\nfine without ever issuing \"git status\" (assuming you're the one\ncontrolling the changes, that is).\n\n\"git rm\" and \"git add\" won't walk the tree. They're just interested in\nthe paths you give them and won't touch anything else.\n\n\"git commit path1 path2\" won't walk the tree. It has to walk the paths\n(which can be entire subdirectories, or all of them), but not more than\nthat.\n\n\"git push\" (ie, send your changes upstream) won't walk the tree. It'll\njust look at the history and how they differ.\n\n\"git merge\" (and therefore also \"git pull\") doesn't walk the tree. It\nonly makes sure paths that are touched by the merge are up-to-date.\n\nApart from that, it would be trivial to hack up some inotify config\nand scripts that stages changes in a separate index-file and then\nadd a simple wrapper that operates on the separate index-file rather\nthan the \"regular\" one.\n\nSample \"giti\" wrapper:\n--%<--%<--%<--\n#!/bin/sh\n# giti - inotify driven git wrapper\nGIT_INDEX=.git/inotify-index\nexport GIT_INDEX\ncase \"$@\" in\n\tstatus)\n\t\tgit diff --name-only --cached\n\t\texit $?\n\t\t;;\nesac\n\ngit \"$@\"\n--%<--%<--%<--\n\nSample inotify script:\n--%<--%<--%<--\n#!/bin/sh\nGIT_INDEX=.git/inotify-index git add $1\n--%<--%<--%<--\n\nSample incrontab(5) entry:\n--%<--%<--%<--\n/watched/path IN_CLOSE_WRITE inotify.git $@/$#\n--%<--%<--%<--\n\nTotally untested ofcourse, so it probably needs tweaking. It should\nwork rather well though, assuming you're somewhat careful what\narguments you send to the \"giti\" wrapper and make sure to never\nuse any git-commands that *have* to walk the entire tree (such as\n\"git commit -a\").\n\nLet us know how it pans out.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"93762","messageId":"49002B27.50201@op5.se","threadId":"16014","inReplyTo":"m3d4hsi708.fsf@localhost.localdomain","subject":"Re: git performance","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-10-23T07:43:35Z","receivedAt":"2008-10-23T07:43:35Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Jakub Narebski wrote:\n> \"Edward Ned Harvey\" <git@nedharvey.com> writes:\n> \n>> I see things all over the Internet saying git is fast.  I'm\n>> currently struggling with poor svn performance and poor attitude of\n>> svn developers, so I'd like to consider switching to git.  A quick\n>> question first.\n>>\n>> The core of the performance problem I'm facing is the need to \"walk\n>> the tree\" for many thousand files.  Every time I do \"svn update\" or\n>> \"svn status\" the svn client must stat every file to check for local\n>> modifications (a coffee cup or a beer worth of stats).  In essence,\n>> this is unavoidable if there is no mechanism to constantly monitor\n>> filesystem activity during normal operations.  Analogous to\n>> filesystem journaling.\n>>\n>> So - I didn't see anything out there saying \"git is fast because it\n>> uses inotify\" or anything like that.  Perhaps git would not help me\n>> at all?  Because git still needs to stat all the files in the tree?\n> \n> http://git.or.cz/gitwiki/GitBenchmarks\n> \n> While it should be possible to use 'assume unchanged' bit together\n> with inotify / icron, it is not something tha is done; IIRC Mercurial\n> had Linux-only InotifyPlugin...\n> \n\nWell, inotify() is Linux specific, so it'd be quite hard to support on\nanother platform. Emulating it with a billion stat() calls feels rather\nlike a disk (and I/O performance) killer.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"93773","messageId":"vpq4p33pkv3.fsf@bauges.imag.fr","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"Re: git performance","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2008-10-23T12:16:32Z","receivedAt":"2008-10-23T12:16:32Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"\"Edward Ned Harvey\" <git@nedharvey.com> writes:\n\n>> Yes, it does stat all the files. How many files are you talking about,\n>> and what platform?  From a warm cache on Linux, the 23,000 files kernel\n>> repo takes about a tenth of a second to stat all files for me (and this\n>> on a several year-old machine). And of course many operations don't\n>> require stat'ing at all (like looking at logs, or diffs that don't\n>> involve the working tree).\n>\n> No worries.  No solution can meet everyone's needs.\n>\n> I'm talking about 40-50,000 files, on multi-user production linux,\n> which means the cache is never warm, except when I'm benchmarking.\n> Specifically RHEL 4 with the files on NFS mount. Cold cache \"svn st\"\n> takes ~10 mins. Warm cache 20-30 sec.\n\nSVN does not only has to stat the files. It also has to read the\nstat-cache information wich is split in one .svn/ per directory in the\nworking tree. Not sure which operation dominates the performance,\nthough. Best is just to try.\n\n> Out of curiosity, what are they talking about, when they say \"git is\n> fast?\" Just the fact that it's all local disk, or is there more to\n> it than that?\n\nNot just local disk: bzr also works locally, and git is much faster on\nmost operations (bzr status can now compete with git, but \"git log\"\nand \"git commit\" can be instantaneous where bzr take 1 minute for\nexample).\n\nFor sure, doing most operations locally is the key to being fast, but\nGit has also been written so that the complexity of algorithms be as\nlow as possible.\n\n> I could see - git would probably outperform perforce for versioning\n> of large files (let's say iso files) to benefit from sustained local\n> disk IO, while perforce would probably outperform anything I can\n> think of, operating on thousands of tiny files, because it will\n> never walk the tree.\n\nMercurial has an extension called \"inotify\" that avoids walking the\ndisk too. AFAIK doesn't have an equivalent in Git (mostly because most\npeople interested find git fast enough).\n\n-- \nMatthieu\n"},{"id":"93778","messageId":"fcaeb9bf0810230604u6db6d31cr91153b3cbfa0bbb6@mail.gmail.com","threadId":"16014","inReplyTo":"49002B27.50201@op5.se","subject":"Re: git performance","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2008-10-23T13:04:07Z","receivedAt":"2008-10-23T13:04:07Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 10/23/08, Andreas Ericsson <ae@op5.se> wrote:\n> Jakub Narebski wrote:\n>\n> > \"Edward Ned Harvey\" <git@nedharvey.com> writes:\n> >\n> >\n> > > I see things all over the Internet saying git is fast.  I'm\n> > > currently struggling with poor svn performance and poor attitude of\n> > > svn developers, so I'd like to consider switching to git.  A quick\n> > > question first.\n> > >\n> > > The core of the performance problem I'm facing is the need to \"walk\n> > > the tree\" for many thousand files.  Every time I do \"svn update\" or\n> > > \"svn status\" the svn client must stat every file to check for local\n> > > modifications (a coffee cup or a beer worth of stats).  In essence,\n> > > this is unavoidable if there is no mechanism to constantly monitor\n> > > filesystem activity during normal operations.  Analogous to\n> > > filesystem journaling.\n> > >\n> > > So - I didn't see anything out there saying \"git is fast because it\n> > > uses inotify\" or anything like that.  Perhaps git would not help me\n> > > at all?  Because git still needs to stat all the files in the tree?\n> > >\n> >\n> > http://git.or.cz/gitwiki/GitBenchmarks\n> >\n> > While it should be possible to use 'assume unchanged' bit together\n> > with inotify / icron, it is not something tha is done; IIRC Mercurial\n> > had Linux-only InotifyPlugin...\n> >\n> >\n>\n>  Well, inotify() is Linux specific, so it'd be quite hard to support on\n>  another platform. Emulating it with a billion stat() calls feels rather\n>  like a disk (and I/O performance) killer.\n\nThere is \"filemon\" on Windows, which monitors file access. I don't\nknow how it impacts performance though. A quick search revealed kqueue\nfor FreeBSD/Mac OSX.\n-- \nDuy\n"},{"id":"93792","messageId":"20081023163912.GA11489@coredump.intra.peff.net","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"Re: git performance","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-10-23T16:39:12Z","receivedAt":"2008-10-23T16:39:12Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Oct 22, 2008 at 05:55:14PM -0400, Edward Ned Harvey wrote:\n\n> I'm talking about 40-50,000 files, on multi-user production linux,\n> which means the cache is never warm, except when I'm benchmarking.\n\nWell, if you have a cold cache it's going to take longer. :) You should\nprobably benchmark if you want to know exactly how long.\n\n> Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn st\"\n> takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me,\n\nWow, that is awful. For comparison, \"git status\" from a cold on the\nkernel repo takes me 17 seconds. From a warm cache, less than half a\nsecond.\n\nYes, the cold cache case would probably be better with inotify, but\ncompared to svn, that's screaming fast. I haven't used perforce. If your\nbottleneck really is stat'ing the tree, then yes, something that avoided\nthat might perform better (but weigh that particular optimization\nagainst other things which might be slower).\n\n> Out of curiosity, what are they talking about, when they say \"git is\n> fast?\"\n\nWell, there are the numbers above. When comparing to SVN or (god forbid)\nCVS, there are order of magnitude speedups for most common operations.\n\n>  Just the fact that it's all local disk, or is there more to it\n> than that?  I could see - git would probably outperform perforce for\n\nThe things that generally make git fast are:\n\n  - using a compact on-disk structure (including zlib and aggressive\n    delta-finding) to keep your cache warm (and when it's not warm, to\n    get data off the disk as quickly as possible)\n\n  - the content-addressable nature of objects means we can just look at\n    the data we need to solve a problem. For example,\n    getting the history between point A and point B is \"O(the number of\n    commits between A and B)\", _not_ \"O(the size of the repo)\".\n    Viewing a log without generating diffs is \"O(the number of\n    commits)\", not \"O(some combination of the number of commits and the\n    number of files in each commit)\". Diffing two points in history is\n    \"O(the size of the differences between the two points)\" and is\n    totally independent of the number of commits between the two points.\n\n  - most operations are streamable. \"git log >/dev/null\" on the kernel\n    repo (about 90,000 commits) takes 8.5 seconds on my box. But it\n    starts generating output immediately, so it _feels_ instant, and the\n    rest of the data is generated while I read the first commit in my\n    pager.\n\n-Peff\n"},{"id":"93797","messageId":"alpine.LNX.1.00.0810231346520.19665@iabervon.org","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"RE: git performance","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-10-23T18:31:16Z","receivedAt":"2008-10-23T18:31:16Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 22 Oct 2008, Edward Ned Harvey wrote:\n\n> Out of curiosity, what are they talking about, when they say \"git is \n> fast?\"  Just the fact that it's all local disk, or is there more to it \n> than that?  I could see - git would probably outperform perforce for \n> versioning of large files (let's say iso files) to benefit from \n> sustained local disk IO, while perforce would probably outperform \n> anything I can think of, operating on thousands of tiny files, because \n> it will never walk the tree. \n\nIt shouldn't be too hard to make git work like perforce with respect to \nwalking the tree. git keeps an index of the stat() info it saw when it \nlast looked at files, and only looks at the contents of files whose stat() \ninfo has changed. In order to have it work like perforce, it would just \nneed to have a flag in the stat() info index for \"don't even bother\", \nwhich it would use for files that aren't \"open\"; for files with this flag, \nthe check for index freshness would always say it's fresh without looking \nat the filesystem. Then you'd just have a config option to check out files \nas \"not open\" (and not writeable), and have a \"git open\" program that \nwould chmod files and get their real stat info.\n\nOf course, git is tuned for cases where the modify/build/test cycle \nrequires stat() (or worse) on every file.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"93807","messageId":"20081024072412.6117@nanako3.lavabit.com","threadId":"16014","inReplyTo":"alpine.LNX.1.00.0810231346520.19665@iabervon.org","subject":"Re: git performance","fromName":"Nanako Shiraishi","fromEmail":"nanako3@lavabit.com","sentAt":"2008-10-23T22:24:12Z","receivedAt":"2008-10-23T22:24:12Z","isPatch":false,"sender":{"key":"nanako3@lavabit.com","avatar":"https://gravatar.com/avatar/3777b9e201c5883a62b1a6fdf7c53f2d712d1d80989146063ea861e33aad72a8?d=mp&s=160"},"body":"Quoting Daniel Barkalow <barkalow@iabervon.org>:\n\n> On Wed, 22 Oct 2008, Edward Ned Harvey wrote:\n>\n>> Out of curiosity, what are they talking about, when they say \"git is \n>> fast?\"  Just the fact that it's all local disk, or is there more to it \n>> than that?  I could see - git would probably outperform perforce for \n>> versioning of large files (let's say iso files) to benefit from \n>> sustained local disk IO, while perforce would probably outperform \n>> anything I can think of, operating on thousands of tiny files, because \n>> it will never walk the tree. \n>\n> It shouldn't be too hard to make git work like perforce with respect to \n> walking the tree. git keeps an index of the stat() info it saw when it \n> last looked at files, and only looks at the contents of files whose stat() \n> info has changed. In order to have it work like perforce, it would just \n> need to have a flag in the stat() info index for \"don't even bother\", \n\nAre you describing the \"assume unchanged bit\"?\n\n-- \nNanako Shiraishi\nhttp://ivory.ap.teacup.com/nanako3/\n"},{"id":"93823","messageId":"alpine.LNX.1.00.0810232237060.19665@iabervon.org","threadId":"16014","inReplyTo":"20081024072412.6117@nanako3.lavabit.com","subject":"Re: git performance","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-10-24T03:56:46Z","receivedAt":"2008-10-24T03:56:46Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 24 Oct 2008, Nanako Shiraishi wrote:\n\n> Quoting Daniel Barkalow <barkalow@iabervon.org>:\n> \n> > On Wed, 22 Oct 2008, Edward Ned Harvey wrote:\n> >\n> >> Out of curiosity, what are they talking about, when they say \"git is \n> >> fast?\"  Just the fact that it's all local disk, or is there more to it \n> >> than that?  I could see - git would probably outperform perforce for \n> >> versioning of large files (let's say iso files) to benefit from \n> >> sustained local disk IO, while perforce would probably outperform \n> >> anything I can think of, operating on thousands of tiny files, because \n> >> it will never walk the tree. \n> >\n> > It shouldn't be too hard to make git work like perforce with respect to \n> > walking the tree. git keeps an index of the stat() info it saw when it \n> > last looked at files, and only looks at the contents of files whose stat() \n> > info has changed. In order to have it work like perforce, it would just \n> > need to have a flag in the stat() info index for \"don't even bother\", \n> \n> Are you describing the \"assume unchanged bit\"?\n\nYes, but with the user write mode bit in the filesystem set to \nno-assume-unchanged, which is how Perforce users cope with it. I hadn't \nrealized it had been implemented to get set on a per-file basis, rather \nthan just as a global setting that caused it to not stat() anything except \nright when it was told to update.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"93835","messageId":"49017F8F.3000908@pcharlan.com","threadId":"16014","inReplyTo":"000901c93490$e0c40ed0$a24c2c70$@com","subject":"Re: git performance","fromName":"Pete Harlan","fromEmail":"pgit@pcharlan.com","sentAt":"2008-10-24T07:55:59Z","receivedAt":"2008-10-24T07:55:59Z","isPatch":false,"sender":{"key":"pgit@pcharlan.com","avatar":null},"body":"Edward Ned Harvey wrote:\n> > Yes, it does stat all the files. How many files are you talking about,\n> > and what platform?  From a warm cache on Linux, the 23,000 files kernel\n> > repo takes about a tenth of a second to stat all files for me (and this\n>\n> I'm talking about 40-50,000 files, on multi-user production linux,\n> which means the cache is never warm, except when I'm benchmarking.\n> Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn\n> st\" takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me,\n\nI did some tests with a repo with ~32k files, and git was slightly\nslower than svn with a cold cache (10.2s vs 8.4s), and around twice as\nfast with a warm cache (.5s vs 1s).\n\nGit 1.6.0.2, svn 1.4.6. Cache made cold with\n\"echo 1 >/proc/sys/vm/drop_caches\".  Timings best of 5 runs.\n\n(I did various benchmarks with svn 1.5.3 also, but there's something\nawfully wrong with svn 1.5.x's merging, which takes pathologically\nlong compared with 1.4 (minutes instead of seconds), and it wasn't\nnoticeably faster than 1.4 at anything I tested.)\n\n> performance was approx the same for files on local disk versus NFS.\n\n10 minutes seems like a crazy amount of time for 40-50k files.  If you\ndidn't say you'd tested it on local disks, it would really sound like\na bad NFS interaction more than an svn problem.\n\n> Out of curiosity, what are they talking about, when they say \"git is\n> fast?\"\n\nIn my comparisons between svn and git, the operation \"checkout\nrevision N of the tree\" (i.e., \"svn update -r 40000\" vs \"git checkout\n302c7476\") took five minutes on subversion and ten seconds using git.\nThe tests were all local, so git wasn't benefiting from being a DVCS,\nit was just eerily fast on some things.  Svn was even that slow when\nthe revisions were 1 commit different, if it was a large enough\ncommit.\n\nI don't check out whole revisions like that very often, but switching\nbetween branches is a similar operation.  It doesn't usually take five\nminutes in svn but it's an interruption, and with git it isn't.\n\nFor almost everything I tried git was faster, but status wasn't really\none of them.  The compelling cases were the number of things that were\nfaster _enough_ to no longer be an interruption, and being a DVCS, and\nrebase, and rebase -i, and gitk, and a smarter blame, and\nbranching/merging support like it's something you'd do all day long,\nnot just when you were forced to.\n\nHTH,\n\n--Pete\n"},{"id":"93858","messageId":"20081024142947.GB11568@coredump.intra.peff.net","threadId":"16014","inReplyTo":"000001c9358f$232bac70$69830550$@com","subject":"Re: git performance","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-10-24T14:29:48Z","receivedAt":"2008-10-24T14:29:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 24, 2008 at 12:15:19AM -0400, Edward Ned Harvey wrote:\n\n> Feel free to forward to the list, if anyone's still talking about it.\n> I already un-subscribed.\n\nPosting is not limited to subscribers, so you can happily continue the\nconversation there by cc'ing the list (and I am cc'ing the list here).\n\n> I did my benchmarking at least two months ago, so I forgot the exact\n> results now, so I ran the benchmark once just now.  I also downloaded\n> git, and did \"git status\" for comparison.  I rebooted the system in\n> between each trial run, to clear the cache.  Here's the results:\n\nSide note: on Linux, it is much easier to clear the cache via\n\n  echo 1 >/proc/sys/vm/drop_caches\n\nthan to reboot for each benchmark.\n\n> Local disk mirror \"time git status\" on the same tree. 17,468 versioned files, so the whole tree is 30,647 including .git files\n> \t0m 25s\tcold cache\n> \t0m 0.2s\twarm cache trial 1\n> \t0m 0.2s\twarm cache trial 2\n\nHmm. That's a lot of increase in files for .git. Did you try repacking\nand then running your test?\n\n> I questioned whether svn and git were causing unnecessary overhead.\n\nSure, they are doing more than just walking. So there is overhead, but\nit's hard to say how much is unnecessary. However, if you were working\nwith an unpacked git, then it may have had to open() a lot of files in\nthe object db (keep in mind that status doesn't just show the difference\nbetween the working tree and the index; it shows the difference between\nthe index and the last commit. So maybe \"git diff\" would be a more\naccurate comparison).\n\n> Conclusions:  \n> * For \"status\" operations on cold cache, large file count, Neither the\n> performance of git or svn approaches the ideal.  Both are an order of\n> magnitude slower than ideal, which is still assuming \"ideal\" requires\n> walking the tree.  A better ideal avoids the need to walk the tree,\n> and has near-zero total cost.\n\nTry your git benchmark again with a packed repo, and I think you will\nfind it approaches the time it takes to walk the tree.\n\nThat being said, if walking the tree is unacceptable to you, then no,\ncurrent git won't work. You would need to patch it to use inotify (once\nupon a time there was some discussion of this, but it never went\nanywhere -- I guess most people work on machines where they can keep the\ncache relatively warm).\n\n-Peff\n"},{"id":"93862","messageId":"dfdaadcd0810241042k1469fc30x62daa19273404edc@mail.gmail.com","threadId":"16014","inReplyTo":"20081024142947.GB11568@coredump.intra.peff.net","subject":"Re: git performance","fromName":"George Shammas","fromEmail":"georgyo@gmail.com","sentAt":"2008-10-24T17:42:00Z","receivedAt":"2008-10-24T17:42:00Z","isPatch":false,"sender":{"key":"georgyo@gmail.com","avatar":"https://gravatar.com/avatar/bd053d9417028b3a61467f80a6d9573899f9a505ee3d776bd7dad0ef9352fe0d?d=mp&s=160"},"body":"If you are really trying to backup a filesystem, you may want to look\nat a filesystem that can do snapshots, it would be a lot more\nefficient then a version control system.  Such as NILFS and ZFS.\n\nhttp://en.wikipedia.org/wiki/NILFS\nhttp://en.wikipedia.org/wiki/ZFS\n\nBoth these will allow you to look at changed files over time. NILFS is\nslightlly diffrent in that it doesn't take snapshots, because it never\ndeletes, so you can rollback every change on a file. They both also\nallow each user to rollback their own files if they wanted to, so if\nthis is your goal, source code version control is not for you, and a\ngood file system is for you.\n\n-G\n\nOn Fri, Oct 24, 2008 at 10:29 AM, Jeff King <peff@peff.net> wrote:\n> On Fri, Oct 24, 2008 at 12:15:19AM -0400, Edward Ned Harvey wrote:\n>\n>> Feel free to forward to the list, if anyone's still talking about it.\n>> I already un-subscribed.\n>\n> Posting is not limited to subscribers, so you can happily continue the\n> conversation there by cc'ing the list (and I am cc'ing the list here).\n>\n>> I did my benchmarking at least two months ago, so I forgot the exact\n>> results now, so I ran the benchmark once just now.  I also downloaded\n>> git, and did \"git status\" for comparison.  I rebooted the system in\n>> between each trial run, to clear the cache.  Here's the results:\n>\n> Side note: on Linux, it is much easier to clear the cache via\n>\n>  echo 1 >/proc/sys/vm/drop_caches\n>\n> than to reboot for each benchmark.\n>\n>> Local disk mirror \"time git status\" on the same tree. 17,468 versioned files, so the whole tree is 30,647 including .git files\n>>       0m 25s  cold cache\n>>       0m 0.2s warm cache trial 1\n>>       0m 0.2s warm cache trial 2\n>\n> Hmm. That's a lot of increase in files for .git. Did you try repacking\n> and then running your test?\n>\n>> I questioned whether svn and git were causing unnecessary overhead.\n>\n> Sure, they are doing more than just walking. So there is overhead, but\n> it's hard to say how much is unnecessary. However, if you were working\n> with an unpacked git, then it may have had to open() a lot of files in\n> the object db (keep in mind that status doesn't just show the difference\n> between the working tree and the index; it shows the difference between\n> the index and the last commit. So maybe \"git diff\" would be a more\n> accurate comparison).\n>\n>> Conclusions:\n>> * For \"status\" operations on cold cache, large file count, Neither the\n>> performance of git or svn approaches the ideal.  Both are an order of\n>> magnitude slower than ideal, which is still assuming \"ideal\" requires\n>> walking the tree.  A better ideal avoids the need to walk the tree,\n>> and has near-zero total cost.\n>\n> Try your git benchmark again with a packed repo, and I think you will\n> find it approaches the time it takes to walk the tree.\n>\n> That being said, if walking the tree is unacceptable to you, then no,\n> current git won't work. You would need to patch it to use inotify (once\n> upon a time there was some discussion of this, but it never went\n> anywhere -- I guess most people work on machines where they can keep the\n> cache relatively warm).\n>\n> -Peff\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"},{"id":"93864","messageId":"alpine.LFD.2.00.0810241050310.3287@nehalem.linux-foundation.org","threadId":"16014","inReplyTo":"20081024142947.GB11568@coredump.intra.peff.net","subject":"Re: git performance","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-10-24T17:53:20Z","receivedAt":"2008-10-24T17:53:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 24 Oct 2008, Jeff King wrote:\n> \n> Side note: on Linux, it is much easier to clear the cache via\n> \n>   echo 1 >/proc/sys/vm/drop_caches\n\nUse \"echo 3\" instead of \"1\".\n\nIt's actually a bitmask, with bit 0 being \"data\" (pagecache) and bit 1 \nbeing \"metadata\" (inodes and directory caches).\n\nAnd since git (or any SCM) is very metadata-intensive, you really should \nmake sure to drop metadata too, otherwise your caches won't be really very \ncold at all.\n\n(But it obviously depends on the operation you're testing - some are more \nabout the inodes and directories, others are about file data access).\n\n\t\t\tLinus\n"},{"id":"93865","messageId":"20081024182031.GA11287@coredump.intra.peff.net","threadId":"16014","inReplyTo":"alpine.LFD.2.00.0810241050310.3287@nehalem.linux-foundation.org","subject":"Re: git performance","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-10-24T18:20:31Z","receivedAt":"2008-10-24T18:20:31Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Oct 24, 2008 at 10:53:20AM -0700, Linus Torvalds wrote:\n\n> >   echo 1 >/proc/sys/vm/drop_caches\n> \n> Use \"echo 3\" instead of \"1\".\n> \n> It's actually a bitmask, with bit 0 being \"data\" (pagecache) and bit 1 \n> being \"metadata\" (inodes and directory caches).\n> \n> And since git (or any SCM) is very metadata-intensive, you really should \n> make sure to drop metadata too, otherwise your caches won't be really very \n> cold at all.\n> \n> (But it obviously depends on the operation you're testing - some are more \n> about the inodes and directories, others are about file data access).\n\nAh, thanks. In this case, he was interested in walking the directory\ntree, so the metadata caching was indeed very important.\n\n-Peff\n"},{"id":"93867","messageId":"m3zlkthkta.fsf@localhost.localdomain","threadId":"16014","inReplyTo":"dfdaadcd0810241042k1469fc30x62daa19273404edc@mail.gmail.com","subject":"Re: git performance","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-10-24T19:06:37Z","receivedAt":"2008-10-24T19:06:37Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"George Shammas\" <georgyo@gmail.com> writes:\n\n> If you are really trying to backup a filesystem, you may want to look\n> at a filesystem that can do snapshots, it would be a lot more\n> efficient then a version control system.  Such as NILFS and ZFS.\n> \n> http://en.wikipedia.org/wiki/NILFS\n> http://en.wikipedia.org/wiki/ZFS\n\nOr ext3cow, or (currently in early stages of development) Tux3\n\n  http://en.wikipedia.org/wiki/Ext3cow\n  http://en.wikipedia.org/wiki/Tux3\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"93899","messageId":"490255D1.8060804@pcharlan.com","threadId":"16014","inReplyTo":"49017F8F.3000908@pcharlan.com","subject":"Re: git performance","fromName":"Pete Harlan","fromEmail":"pgit@pcharlan.com","sentAt":"2008-10-24T23:10:09Z","receivedAt":"2008-10-24T23:10:09Z","isPatch":false,"sender":{"key":"pgit@pcharlan.com","avatar":null},"body":"Pete Harlan wrote:\n> Edward Ned Harvey wrote:\n>>> Yes, it does stat all the files. How many files are you talking about,\n>>> and what platform?  From a warm cache on Linux, the 23,000 files kernel\n>>> repo takes about a tenth of a second to stat all files for me (and this\n>> I'm talking about 40-50,000 files, on multi-user production linux,\n>> which means the cache is never warm, except when I'm benchmarking.\n>> Specifically RHEL 4 with the files on NFS mount.  Cold cache \"svn\n>> st\" takes ~10 mins.  Warm cache 20-30 sec.  Surprisingly to me,\n> \n> I did some tests with a repo with ~32k files, and git was slightly\n> slower than svn with a cold cache (10.2s vs 8.4s), and around twice as\n> fast with a warm cache (.5s vs 1s).\n> \n> Git 1.6.0.2, svn 1.4.6. Cache made cold with\n> \"echo 1 >/proc/sys/vm/drop_caches\".  Timings best of 5 runs.\n\nAfter redoing this test with \"echo 3 >/proc/sys/vm/drop_caches\" (which\nalso discards metadata, as pointed out by Linus), the cold-cache\ntimings are:\n\n\tsvn 12.65 seconds\n\tgit 10.3  seconds\n\nSo no Earth-shattering difference, but now git is somewhat quicker\nthan Subversion at everything I tested.\n\n--Pete\n"}]}