{"thread":{"id":"781","subject":"Mercurial 0.5b vs git","startedAt":"2005-05-31T21:31:03Z","lastAt":"2005-05-31T21:31:03Z","messageCount":1,"participants":["Matt Mackall"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"4320","messageId":"20050531213103.GR7685@waste.org","threadId":"781","inReplyTo":null,"subject":"Mercurial 0.5b vs git","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-31T21:31:03Z","receivedAt":"2005-05-31T21:31:03Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"The latest version of Mercurial is available at:\n\n http://selenic.com/mercurial/\n\nUtilities to convert git repos and interoperate with git are beginning\nto appear on the mercurial mailing list, including a port of gitk.\n\nAs a practical demonstration, I've imported Ingo's BKCVS patchset into\nMercurial. The result is a 297M archive with 28237 changesets going back\nto 2.4.0. Some history is lost because of the BK->CVS flattening. You\ncan browse it here:\n\n http://userweb.kernel.org/~mpm/linux-hg/index.cgi\n\nBe sure to check out the annotate feature. Unfortunately there are no\nbranches in this repo because of the BK->CVS flattening, but you can\nlook at the main Mercurial repo to see examples of pulls.\n\nThe full tarball of the Mercurial kernel repo (144MB) can be grabbed here:\n\n http://www.kernel.org/pub/linux/kernel/people/mpm/linux-hg.tar.gz\n\nIf you want to browse this repo on your own machine (very fast and\nconvenient for laptops!), simply install Mercurial, download the\ntarball, run 'hg serve' in the repo directory and point your web\nbrowser at http://localhost:8000.\n\nThe web interface also serves as a highly efficient merge server:\n\n$ time hg -v merge http://remotehost:8000/\nsearching for changes\nadding changesets\nadding manifests\nadding files\n118549846 bytes of data transfered\nmodified 23306 files, added 28238 changesets and 188476 new revisions\n\nreal    4m51.371s\nuser    1m25.852s\nsys     0m8.303s\n\nThat's pulling the whole kernel history over fast DSL with only 113M\nof traffic. Compare that to the 2.6.11 tar.bz2 at 35M. Smaller merges\nare of course proportionally faster. (Pulls from userweb.kernel.org\nare disabled because the machine has limited bandwidth.)\n\nVerifying the archive:\n\n$ time hg verify\nchecking changesets\nchecking manifests\ncrosschecking files in changesets and manifests\nchecking files\n23305 files, 28238 changesets, 188464 total revisions\n\nreal    2m48.986s\nuser    1m30.055s\nsys     0m7.158s\n\nChecking the integrity of the equivalent git archive looks like it\nwill take an hour or more of seek intensive I/O (though the person\nwho was timing it for me gave up).\n\nThis highlights one of git's most serious problems: storing the\nrepository by hash. This tends to pessimize layout over time. Initial\ncheck-ins will be nicely ordered by write order, but as changes are\nmade, the set of files in the tip will get spread further and further\napart on the disk and in more and more random order. Copying the\narchive via rsync, cp -a, or the like will tend to exacerbate things\nby reordering _everything_ in hash (aka worst possible) order. This is\npretty fundamental to the git design and will cause its scalability to\nfall apart as the number of revisions mount.\n\nMercurial was originally using a similar scheme, and when I ran into\nthis problem, I spent a day playing with variations on sorting by\ninode, prefetching, etc to get the performance back. None of it came\nclose to the performance of simply having everything layed out well on\ndisk in the first place.\n\nMy eventual solution was a simple 5-line change to switch back to a\ntree-structured repo layout like CVS. This lets the filesystem block\nallocator assist by putting files in the same directory near each\nother on disk. Also, copying repos tends to optimize things rather\nthan making things worse. Mercurial also inherently stores all file\nrevisions together so operations like tree diffs or file annotate can\nbe done with a minimum of seeking.\n\n\nHere's a quick comparison:\n\n                    Mercurial      git                     BK (*)\nstorage             revlog delta   compressed revisions    SCCS weave\nstorage naming      by filename    by revision hash        by filename\nmerge               file DAGs      changeset DAG           file DAGs?\nconsistency         SHA1           SHA1                    CRC\nsignable?           yes            yes                     no       \n\nretrieve file tip   O(1)           O(1)                    O(revs)\nadd rev             O(1)           O(1)                    O(revs)\nfind prev file rev  O(1)           O(changesets)           O(revs)\nannotate file       O(revs)        O(changesets)           O(revs)\nfind file changeset O(1)           O(changesets)           ?\n\nfile tracking       stat-based     stat-based              bk edit\ncheckout            O(files)       O(files)                O(revs)?\ncommit              O(changes)     O(changes)              ?\n                    6 patches/s    6 patches/s             slow\ndiff working dir    O(changes)     O(changes)              ?\n                    < 1s           < 1s                    ?\ntree diff revs      O(changes)     O(changes)              ?\n                    < 1s           < 1s                    ?\nhardlink clone      O(files)       O(revisions)            O(files)\n\nfind remote csets   O(log new)     rsync: O(revisions)     ?\n                                   git-http: O(changesets)\npull remote csets   O(patch)       O(modified files)       O(patch)\n\nrepo growth         O(patch)       O(revisions)            O(patch)\n kernel history     297M           3.5G?                   250M?\nlines of code       3700           6500+cogito+gitweb+..   ??\n\n* I've never used BK so this is just guesses\n\n-- \nMathematics is the supreme nostalgia of our time.\n\n"}]}