{"thread":{"id":"36107","subject":"question about: Facebook makes Mercurial faster than Git","startedAt":"2014-03-10T10:07:21Z","lastAt":"2014-03-14T12:58:15Z","messageCount":12,"participants":["Dennis Luehring","David Lang","demerphq","Johan Herland","Karsten Blees","Michael Haggerty","Ondřej Bílka","Martin Langhoff","Duy Nguyen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"236367","messageId":"531D8ED9.7040305@gmx.net","threadId":"36107","inReplyTo":null,"subject":"question about: Facebook makes Mercurial faster than Git","fromName":"Dennis Luehring","fromEmail":"dl.soluz@gmx.net","sentAt":"2014-03-10T10:07:21Z","receivedAt":"2014-03-10T10:07:21Z","isPatch":false,"sender":{"key":"dl.soluz@gmx.net","avatar":null},"body":"according to these blog posts\n\nhttp://www.infoq.com/news/2014/01/facebook-scaling-hg\nhttps://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/\n\nmercurial \"can\" be faster then git\n\nbut i don't found any reply from the git community if it is a real problem\nor if there a ongoing (maybe git 2.0) changes to compete better in this case\n"},{"id":"236368","messageId":"alpine.DEB.2.02.1403100310080.25193@nftneq.ynat.uz","threadId":"36107","inReplyTo":"531D8ED9.7040305@gmx.net","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2014-03-10T10:13:45Z","receivedAt":"2014-03-10T10:13:45Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 10 Mar 2014, Dennis Luehring wrote:\n\n> according to these blog posts\n>\n> http://www.infoq.com/news/2014/01/facebook-scaling-hg\n> https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/\n>\n> mercurial \"can\" be faster then git\n>\n> but i don't found any reply from the git community if it is a real problem\n> or if there a ongoing (maybe git 2.0) changes to compete better in this case\n\nAs I understand this, the biggest part of what happened is that Facebook made a \ntweak to mercurial so that when it needs to know what files have changed in \ntheir massive tree, their version asks their special storage array, while git \nwould have to look at it through the filesystem interface (by doing stat calls \non the directories and files to see if anything has changed)\n\nIn other words, unless you have a very high end storage system that can keep \ntrack of such things for you, the Facebook 'fix' won't help you. And even if it \ndoes have such a capability, unless you use the same storage system that \nFacebook uses, you would have to port it to your class of device.\n\nNow, in addition to this, they did some other tweaks and changes, but compared \nto this status change, everything else is minor.\n\nDavid Lang\n"},{"id":"236371","messageId":"CANgJU+W+f3KUxehDGxd+f77RO24VadsnOV=szE2MkBXjs8wDCQ@mail.gmail.com","threadId":"36107","inReplyTo":"531D8ED9.7040305@gmx.net","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2014-03-10T11:28:37Z","receivedAt":"2014-03-10T11:28:37Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On 10 March 2014 11:07, Dennis Luehring <dl.soluz@gmx.net> wrote:\n> according to these blog posts\n>\n> http://www.infoq.com/news/2014/01/facebook-scaling-hg\n> https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/\n>\n> mercurial \"can\" be faster then git\n>\n> but i don't found any reply from the git community if it is a real problem\n> or if there a ongoing (maybe git 2.0) changes to compete better in this case\n\nThey mailed the list about performance issues in git. From what I saw\nthere was relatively little feedback.\n\nI had the impression, and I would not be surprised if they had the\nimpression that the git development community is relatively\nunconcerned about performance issues on larger repositories.\n\nThere have been other reports, which are difficult to keep track of\nwithout a bug tracking system, but the ones I know of are:\n\nPoor performance of git status with large number of excluded files and\nlarge repositories.\nPoor performance, and breakage, on repositories with very large\nnumbers of files in them. (Rebase for instance will break if you\nrebase a commit that contains a *lot* of files.)\nPoor performance in protocol layer (and other places) with repos with\nlarge numbers of refs. (Maybe this is fixed, not sure.)\n\ncheers,\nYves\n\n\n\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"236372","messageId":"531DA519.8090509@gmx.net","threadId":"36107","inReplyTo":"CANgJU+W+f3KUxehDGxd+f77RO24VadsnOV=szE2MkBXjs8wDCQ@mail.gmail.com","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Dennis Luehring","fromEmail":"dl.soluz@gmx.net","sentAt":"2014-03-10T11:42:17Z","receivedAt":"2014-03-10T11:42:17Z","isPatch":false,"sender":{"key":"dl.soluz@gmx.net","avatar":null},"body":"Am 10.03.2014 12:28, schrieb demerphq:\n> I had the impression, and I would not be surprised if they had the\n> impression that the git development community is relatively\n> unconcerned about performance issues on larger repositories.\n\nso the question is if the git community is interested in beeing \ncompetive in such\nlarge scale scenarios - something what mercurial seems to be now out of \nthe box\n"},{"id":"236374","messageId":"CALKQrgcfTKy0d_BGAZ-bSx5i-=MVEF-WuRfW6T3Q-YxvVSqY_A@mail.gmail.com","threadId":"36107","inReplyTo":"531DA519.8090509@gmx.net","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2014-03-10T12:10:07Z","receivedAt":"2014-03-10T12:10:07Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Mon, Mar 10, 2014 at 12:42 PM, Dennis Luehring <dl.soluz@gmx.net> wrote:\n> Am 10.03.2014 12:28, schrieb demerphq:\n>\n>> I had the impression, and I would not be surprised if they had the\n>> impression that the git development community is relatively\n>> unconcerned about performance issues on larger repositories.\n>\n> so the question is if the git community is interested in beeing competive in\n> such large scale scenarios - something what mercurial seems to be now out\n> of the box\n\nAFAIK, David Lang's comment is not far off the mark. Facebook has made\na tool called Watchman (https://github.com/facebook/watchman) that\nwatches your work tree (i.e. wrapping inotify on Linux) and triggers\nvarious commands when files within are changed (e.g. do an auto-build\nwhenever a file in your project changes). Since this tool will\ndiscover when files change, they have adjusted Mercurial to discover\nchanges by querying Watchman instead of stat-ing the entire work tree.\n\nAFAICS, this is basically a tradeoff between the time it takes to stat\nyour work tree and the overhead/administrivia of running a daemon to\nmonitor the work tree. It seems Facebook has organized their code and\ninfrastructure in a way that makes the latter approach worthwhile for\nthem, and has contributed their solution back to Mercurial.\n\nIt should be possible to teach Git to do similar things, and IINM\nthere are (and have previously been) several attempts to do similar\nthings in Git, e.g.:\n\n - http://thread.gmane.org/gmane.comp.version-control.git/240339\n\n - http://thread.gmane.org/gmane.comp.version-control.git/217817\n\nI haven't looked closely at these attempts (it is not my scratch to\nitch), and I don't know if/how they would work on top of Watchman, but\nin principle I don't see why Git shouldn't be able to leverage\nWatchman the same way Mercurial does.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"236405","messageId":"531DC9B6.6030605@gmail.com","threadId":"36107","inReplyTo":"531DA519.8090509@gmx.net","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2014-03-10T14:18:30Z","receivedAt":"2014-03-10T14:18:30Z","isPatch":false,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 10.03.2014 12:42, schrieb Dennis Luehring:\n> Am 10.03.2014 12:28, schrieb demerphq:\n>> I had the impression, and I would not be surprised if they had the\n>> impression that the git development community is relatively\n>> unconcerned about performance issues on larger repositories.\n> \n> so the question is if the git community is interested in beeing competive in such\n> large scale scenarios - something what mercurial seems to be now out of the box\n> \n\nThe hgwatchman site claims (https://bitbucket.org/facebook/hgwatchman)\n\n\"On a real-world repository with over 200,000 files, hg status normally takes over 3 seconds. With hgwatchman it takes under 0.6 seconds.\"\n\nThere have been a few performance improvements in git status to support such large repositories. I just re-checked git status performance with the WebKit repo (~200k files):\n\nLinux (with core.preloadIndex)\ngit status -uall: 0.620s\ngit status -uno : 0.255s\n\nWindows (with core.preloadIndex and core.fscache)\ngit status -uall: 1.006s\ngit status -uno : 0.695s\n\nOf course, for more reliable benchmark data, you'd have to compare the same repo on the same platform. But on first glance, it seems that mercurial with hgwatchman extension may be as fast as git is out of the box, not the other way around.\n\nThis comes at the cost of running a background daemon, which may slow down the entire system. E.g. if the daemon activates whenever the compiler creates a .o file, it will probably slow down build performance.\n\nNote that hgwatchman doesn't support Windows, so git is probably much faster there.\n"},{"id":"236407","messageId":"531DD0B4.3020603@alum.mit.edu","threadId":"36107","inReplyTo":"CALKQrgcfTKy0d_BGAZ-bSx5i-=MVEF-WuRfW6T3Q-YxvVSqY_A@mail.gmail.com","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2014-03-10T14:48:20Z","receivedAt":"2014-03-10T14:48:20Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 03/10/2014 01:10 PM, Johan Herland wrote:\n> It should be possible to teach Git to do similar things, and IINM\n> there are (and have previously been) several attempts to do similar\n> things in Git, e.g.:\n> \n>  - http://thread.gmane.org/gmane.comp.version-control.git/240339\n> \n>  - http://thread.gmane.org/gmane.comp.version-control.git/217817\n> \n> I haven't looked closely at these attempts (it is not my scratch to\n> itch), and I don't know if/how they would work on top of Watchman, but\n> in principle I don't see why Git shouldn't be able to leverage\n> Watchman the same way Mercurial does.\n\nThis touches on the most important thing that we should take to heart\nfrom this episode:\n\nOf course Facebook could have modified either Git or Mercurial to do\nwhat they want.  Why did they pick Mercurial?  The article seems to\nclaim that they were initially biased towards Git, but they chose\nMercurial because its code base is easier to modify.  This is a claim\nthat I can easily believe.\n\nThe two projects are almost exactly the same age.  The number of commits\nin the two projects is similar.  Mercurial has had fewer contributors\nactive at any given time over its project lifetime.\n\nBut let's see how much code is in the main part of Mercurial vs. Git:\n\n    $ find mercurial hgext \\( -name '*.c' -o -name '*.py' \\) -print |\n          xargs cat | wc -l\n    46164\n\n    $ cat *.c *.h *.sh *.perl builtin/*.c | wc -l\n    188530\n\nThese are just crude estimates and I hope I got the right directories\nfor Mercurial.  But, by these numbers, Git has 4 times as much code as\nMercurial.  That alone will go a long way to making Git harder to\nmodify.  I don't think that Git has anywhere near 4 times the features\nof Mercurial.  Probably most of the difference can be explained by the\nchoice of implementation languages; 94% of the code in these hg\ndirectories is Python, whereas 88% of Git's core code is C.\n\nHow can we make Git easier to hack (short of switching languages)?  Here\nare my suggestions:\n\n* Better function docstrings -- don't make developers have to read the\nwhole call stack to find out what a function does, or who owns the\nmemory that is passed around.\n\n* More modularity -- more coherent and abstract APIs between different\nparts of the system, and less pawing around in your neighbor's data\nstructures.\n\n* Higher-level abstractions -- make more use of APIs like strbuf and\nstring_list as opposed to handling every malloc() and realloc() by hand.\n\nI personally wish that we as a project would be more willing to spend a\nfew extra CPU microseconds to make our code easier to read and modify\nand more robust.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"236427","messageId":"20140310175102.GA17336@domone.podge","threadId":"36107","inReplyTo":"alpine.DEB.2.02.1403100310080.25193@nftneq.ynat.uz","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Ondřej Bílka","fromEmail":"neleai@seznam.cz","sentAt":"2014-03-10T17:51:02Z","receivedAt":"2014-03-10T17:51:02Z","isPatch":false,"sender":{"key":"neleai@seznam.cz","avatar":"https://avatars.githubusercontent.com/u/48067?v=4"},"body":"On Mon, Mar 10, 2014 at 03:13:45AM -0700, David Lang wrote:\n> On Mon, 10 Mar 2014, Dennis Luehring wrote:\n> \n> >according to these blog posts\n> >\n> >http://www.infoq.com/news/2014/01/facebook-scaling-hg\n> >https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/\n> >\n> >mercurial \"can\" be faster then git\n> >\n> >but i don't found any reply from the git community if it is a real problem\n> >or if there a ongoing (maybe git 2.0) changes to compete better in this case\n> \n> As I understand this, the biggest part of what happened is that\n> Facebook made a tweak to mercurial so that when it needs to know\n> what files have changed in their massive tree, their version asks\n> their special storage array, while git would have to look at it\n> through the filesystem interface (by doing stat calls on the\n> directories and files to see if anything has changed)\n> \nThat is mostly a kernel problem. Long ago there was proposed patch to\nadd a recursive mtime so you could check what subtrees changed. If\nsomebody ressurected that patch it would gave similar boost.\n\nThere are two issues that need to be handled, first if you are concerned\nabout one mtime change doing lot of updates a application needs to mark\nall directories it is interested on, when we do update we unmark\ndirectory and by that we update each directory at most once per\napplication run.\n\nSecond problem were hard links where probably a best course is keep list\nof these and stat them separately.\n"},{"id":"236426","messageId":"alpine.DEB.2.02.1403101053120.20306@nftneq.ynat.uz","threadId":"36107","inReplyTo":"20140310175102.GA17336@domone.podge","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2014-03-10T17:56:51Z","receivedAt":"2014-03-10T17:56:51Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 10 Mar 2014, Ondřej Bílka wrote:\n\n> On Mon, Mar 10, 2014 at 03:13:45AM -0700, David Lang wrote:\n>> On Mon, 10 Mar 2014, Dennis Luehring wrote:\n>>\n>>> according to these blog posts\n>>>\n>>> http://www.infoq.com/news/2014/01/facebook-scaling-hg\n>>> https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/\n>>>\n>>> mercurial \"can\" be faster then git\n>>>\n>>> but i don't found any reply from the git community if it is a real problem\n>>> or if there a ongoing (maybe git 2.0) changes to compete better in this case\n>>\n>> As I understand this, the biggest part of what happened is that\n>> Facebook made a tweak to mercurial so that when it needs to know\n>> what files have changed in their massive tree, their version asks\n>> their special storage array, while git would have to look at it\n>> through the filesystem interface (by doing stat calls on the\n>> directories and files to see if anything has changed)\n>>\n> That is mostly a kernel problem. Long ago there was proposed patch to\n> add a recursive mtime so you could check what subtrees changed. If\n> somebody ressurected that patch it would gave similar boost.\n\nbtrfs could actually implement this efficiently, but for a lot of other \nfilesysems this could be very expensive. The question is if it could be enough \nof a win to make it a good choice for people who are doing a heavy git workload \nas opposed to more generic uses.\n\nthere's also the issue of managed vs generated files, if you update the mtime \nall the way up the tree because a source file was compiled and a binary created, \nthat will quickly defeat the value of the recursive mtime.\n\nDavid Lang\n\n> There are two issues that need to be handled, first if you are concerned\n> about one mtime change doing lot of updates a application needs to mark\n> all directories it is interested on, when we do update we unmark\n> directory and by that we update each directory at most once per\n> application run.\n>\n> Second problem were hard links where probably a best course is keep list\n> of these and stat them separately."},{"id":"236453","messageId":"CACPiFC+yjwakzC-0Z=Asuy6SJAxg=pHv4mis_AP_qKVHkpk1Ag@mail.gmail.com","threadId":"36107","inReplyTo":"alpine.DEB.2.02.1403101053120.20306@nftneq.ynat.uz","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2014-03-10T20:22:08Z","receivedAt":"2014-03-10T20:22:08Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Mon, Mar 10, 2014 at 1:56 PM, David Lang <david@lang.hm> wrote:\n> there's also the issue of managed vs generated files, if you update the\n> mtime all the way up the tree because a source file was compiled and a\n> binary created, that will quickly defeat the value of the recursive mime.\n\nI think this points us again to an inotify-based strategy, where git\ncan put an event listener daemon which registers just the watchers it\nneeds, and filters the events on its own conditions.\n\nThe kernel and fs have no good way of knowing about this stuff.\n\ncheers,\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"236505","messageId":"20140311142325.GB17336@domone.podge","threadId":"36107","inReplyTo":"alpine.DEB.2.02.1403101053120.20306@nftneq.ynat.uz","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Ondřej Bílka","fromEmail":"neleai@seznam.cz","sentAt":"2014-03-11T14:23:25Z","receivedAt":"2014-03-11T14:23:25Z","isPatch":false,"sender":{"key":"neleai@seznam.cz","avatar":"https://avatars.githubusercontent.com/u/48067?v=4"},"body":"On Mon, Mar 10, 2014 at 10:56:51AM -0700, David Lang wrote:\n> On Mon, 10 Mar 2014, Ondřej Bílka wrote:\n> \n> >On Mon, Mar 10, 2014 at 03:13:45AM -0700, David Lang wrote:\n> >>On Mon, 10 Mar 2014, Dennis Luehring wrote:\n> >>\n> >>>according to these blog posts\n> >>>\n> >>>http://www.infoq.com/news/2014/01/facebook-scaling-hg\n> >>>https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/\n> >>>\n> >>>mercurial \"can\" be faster then git\n> >>>\n> >>>but i don't found any reply from the git community if it is a real problem\n> >>>or if there a ongoing (maybe git 2.0) changes to compete better in this case\n> >>\n> >>As I understand this, the biggest part of what happened is that\n> >>Facebook made a tweak to mercurial so that when it needs to know\n> >>what files have changed in their massive tree, their version asks\n> >>their special storage array, while git would have to look at it\n> >>through the filesystem interface (by doing stat calls on the\n> >>directories and files to see if anything has changed)\n> >>\n> >That is mostly a kernel problem. Long ago there was proposed patch to\n> >add a recursive mtime so you could check what subtrees changed. If\n> >somebody ressurected that patch it would gave similar boost.\n> \n> btrfs could actually implement this efficiently, but for a lot of\n> other filesysems this could be very expensive. The question is if it\n> could be enough of a win to make it a good choice for people who are\n> doing a heavy git workload as opposed to more generic uses.\n>\nRead next paragraph how do that efficiently, a directory update needs to be done\nonly between application runs. Also there is no overhead when not used\n(except if that makes headers bigger.)\n \n> there's also the issue of managed vs generated files, if you update\n> the mtime all the way up the tree because a source file was compiled\n> and a binary created, that will quickly defeat the value of the\n> recursive mtime.\n>\nYou could do marking on per-file basis. I am not sure if that is needed\nas larger projects use makefiles to not recompile everything so its\nprobably recompiled because source at same directory changed. Also if\nyour compile time is five minutes a half second status would not make\nmuch difference.\n\n \n> \n> >There are two issues that need to be handled, first if you are concerned\n> >about one mtime change doing lot of updates a application needs to mark\n> >all directories it is interested on, when we do update we unmark\n> >directory and by that we update each directory at most once per\n> >application run.\n> >\n> >Second problem were hard links where probably a best course is keep list\n> >of these and stat them separately.\n"},{"id":"236725","messageId":"CACsJy8CP57WqQ1k3jhqZpypua0RimJbE2K5K=WCyheDMk=5L+g@mail.gmail.com","threadId":"36107","inReplyTo":"CANgJU+W+f3KUxehDGxd+f77RO24VadsnOV=szE2MkBXjs8wDCQ@mail.gmail.com","subject":"Re: question about: Facebook makes Mercurial faster than Git","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2014-03-14T12:58:15Z","receivedAt":"2014-03-14T12:58:15Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Mar 10, 2014 at 6:28 PM, demerphq <demerphq@gmail.com> wrote:\n> I had the impression, and I would not be surprised if they had the\n> impression that the git development community is relatively\n> unconcerned about performance issues on larger repositories.\n>\n> There have been other reports, which are difficult to keep track of\n> without a bug tracking system, but the ones I know of are:\n>\n> Poor performance of git status with large number of excluded files and\n> large repositories.\n\nI thought this has been improved lately.. I think we could do better\nstill, but my wip is nowhere ready for anybody's eyes.\n\n> Poor performance, and breakage, on repositories with very large\n> numbers of files in them.\n\nindex v5 and sparse checkout should help a bit. The ultimate solution,\nthough, is narrow clone that's nowhere near finishing. Well, if you\nneed all files present in worktree, then narrow clone does not help\neither..\n\nOn the same line, poor performance on repos with a lot of very large\nfiles also. Junio's split-blob series was a start, but no one picked\nit up, so I guess your impression was right.\n\n> (Rebase for instance will break if you rebase a commit that contains a *lot* of files.)\n\nInteresting. I guess it hits shell's limitations? Roughly how many\nfiles to break it?\n\n> Poor performance in protocol layer (and other places) with repos with\n> large numbers of refs. (Maybe this is fixed, not sure.)\n\nAh.. no it's not. It's being stirred up again though, in both protocol\nand ref backend.\n-- \nDuy\n"}]}