git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: question about: Facebook makes Mercurial faster than Git

From
Ondřej Bílka <neleai@seznam.cz>
Date
Mar 11, 2014, 14:23 UTC
Message-ID
<20140311142325.GB17336@domone.podge>
In-Reply-To
<alpine.DEB.2.02.1403101053120.20306@nftneq.ynat.uz>
On Mon, Mar 10, 2014 at 10:56:51AM -0700, David Lang wrote:
Show 31 quoted lines
> On Mon, 10 Mar 2014, Ondřej Bílka wrote:
> 
> >On Mon, Mar 10, 2014 at 03:13:45AM -0700, David Lang wrote:
> >>On Mon, 10 Mar 2014, Dennis Luehring wrote:
> >>
> >>>according to these blog posts
> >>>
> >>>http://www.infoq.com/news/2014/01/facebook-scaling-hg
> >>>https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/
> >>>
> >>>mercurial "can" be faster then git
> >>>
> >>>but i don't found any reply from the git community if it is a real problem
> >>>or if there a ongoing (maybe git 2.0) changes to compete better in this case
> >>
> >>As I understand this, the biggest part of what happened is that
> >>Facebook made a tweak to mercurial so that when it needs to know
> >>what files have changed in their massive tree, their version asks
> >>their special storage array, while git would have to look at it
> >>through the filesystem interface (by doing stat calls on the
> >>directories and files to see if anything has changed)
> >>
> >That is mostly a kernel problem. Long ago there was proposed patch to
> >add a recursive mtime so you could check what subtrees changed. If
> >somebody ressurected that patch it would gave similar boost.
> 
> btrfs could actually implement this efficiently, but for a lot of
> other filesysems this could be very expensive. The question is if it
> could be enough of a win to make it a good choice for people who are
> doing a heavy git workload as opposed to more generic uses.
>
Read next paragraph how do that efficiently, a directory update needs to be done
only between application runs. Also there is no overhead when not used
(except if that makes headers bigger.)
 
Show 5 quoted lines
> there's also the issue of managed vs generated files, if you update
> the mtime all the way up the tree because a source file was compiled
> and a binary created, that will quickly defeat the value of the
> recursive mtime.
>

You could do marking on per-file basis. I am not sure if that is needed as larger projects use makefiles to not recompile everything so its probably recompiled because source at same directory changed. Also if your compile time is five minutes a half second status would not make much difference.

 
Show 9 quoted lines
> 
> >There are two issues that need to be handled, first if you are concerned
> >about one mtime change doing lot of updates a application needs to mark
> >all directories it is interested on, when we do update we unmark
> >directory and by that we update each directory at most once per
> >application run.
> >
> >Second problem were hard links where probably a best course is keep list
> >of these and stat them separately.
Previous: Martin LanghoffNext: demerphq
Message 6 of 12 in “question about: Facebook makes Mercurial faster than Git”
  1. Dennis LuehringMar 10, 2014
  2. David LangMar 10, 2014
  3. Ondřej BílkaMar 10, 2014
  4. David LangMar 10, 2014
  5. Martin LanghoffMar 10, 2014
  6. Ondřej BílkaMar 11, 2014
  7. demerphqMar 10, 2014
  8. Dennis LuehringMar 10, 2014
  9. Johan HerlandMar 10, 2014
  10. Michael HaggertyMar 10, 2014
  11. Karsten BleesMar 10, 2014
  12. Duy NguyenMar 14, 2014

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.