{"thread":{"id":"410","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","startedAt":"2005-04-30T14:44:17Z","lastAt":"2005-04-30T16:06:40Z","messageCount":2,"participants":["Adam J. Richter","Matt Mackall"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"2260","messageId":"200504301444.j3UEiHN05686@adam.yggdrasil.com","threadId":"410","inReplyTo":null,"subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Adam J. Richter","fromEmail":"adam@yggdrasil.com","sentAt":"2005-04-30T14:44:17Z","receivedAt":"2005-04-30T14:44:17Z","isPatch":false,"sender":{"key":"adam@yggdrasil.com","avatar":null},"body":"On 2005-04-30, Andrea Arcangeli wrote:\n>On a bit more technical side, one thing I'm wondering about is the\n>compression. If I change mercurial like this:\n>\n>--- revlog.py.~1~       2005-04-29 01:33:14.000000000 +0200\n>+++ revlog.py   2005-04-30 03:54:12.000000000 +0200\n>@@ -11,9 +11,11 @@\n> import zlib, struct, mdiff, sha, binascii, os, tempfile\n> \n> def compress(text):\n>+    return text\n>     return zlib.compress(text)\n> \n> def decompress(bin):\n>+    return text\n>     return zlib.decompress(bin)\n> \n> def hash(text):\n>\n>\n>the .hg directory sizes changes from 167M to 302M _BUT_ the _compressed_\n>size of the .hg directory (i.e. like in a full network transfer with\n>rsync -z or a tar.gz backup) changes from 55M to 38M:\n>\n>andrea@opteron:~/devel/kernel> du -sm hg-orig hg-aa hg-orig.tar.bz2 hg-aa.tar.bz2 \n>167     hg-orig\n>302     hg-aa\n>55      hg-orig.tar.bz2\n>38      hg-aa.tar.bz2\n>^^^^^^^^^^^^^^^^^^^^^ 38M backup and network transfer is what I want\n>\n>So I don't really see an huge benefit in compression, other than to\n>slowdown the checkins measurably [i.e. what Linus doesn't want] (the\n>time of compression is a lot higher than the time of python runtime during\n>checkin, so it's hard to believe your 100% boost with psyco in the hg file,\n>sometime psyco doesn't make any difference infact, I'd rather prefer people to\n>work on the real thing of generating native bytecode at compile time, rather\n>than at runtime, like some haskell compiler can do).\n>\n>mercurial is already good at decreasing the entropy by using an efficient\n>storage format, it doesn't need to cheat by putting compression on each blob\n>that can only leads to bad ratios when doing backups and while transferring\n>more than one blob through the network.\n\n\tI'd like to mention a couple of possible optimizations\nfor both the with and without compression approaches.\n\n\tIf you remove the gzip compression, then I imagine you could\ndo much of the IO of checking out files via sendfile, without\never copying data to program space or even changing the program's\nmemory map.  There apparently exists a python sendfile module.\n\n\tIf this mercurial were written in C, much of the rest of\nthe IO could be optimized with mmap (to reduce copies) and writev\nin the absense of a compression pass.  I don't know enough about\npython to know if these optimizations are available.\n\n\tOn the other hand, if you recognize that there is a\nduplication of the work of matching common substrings in\nattepmting to store files as differences and in most compression\nalgorithms, including zlib or bzip2, then you might want to\nconsider storing the files in a format like zdelta or vcdiff, where\ndifferential storage and compression are combined by describing\na file in terms of copy operations both from other files and\n_earlier byte ranges of itself_.\n\n\tzdelta is a modification of zlib for this purpose, but\nI see no permission grants associated with the author's copyright,\nand I thought that zlib only looked at the previous 32kB of data.\n\n\tAlso, if you go this route, you might want to skip the\nlast phases of these compressors where they convert individual\ncharacters into a more compact representation, which I think\nwould defeat inter-file pattern matching if you try to make\na compressed tar of the repository, and would preclude the\nsendfile/mmap optimization (although they might not be worth\nit at this level of granularity).  Then again, since you're\nnaming your files by sha1 hashes, it follows that related files\nwill not be farther apart as the repository grows, so the\ncompression opportunities for larger repositories might be\nless anyhow.\n\n                    __     ______________\nAdam J. Richter        \\ /\nadam@yggdrasil.com      | g g d r a s i l\n"},{"id":"2261","messageId":"20050430160640.GK21897@waste.org","threadId":"410","inReplyTo":"200504301444.j3UEiHN05686@adam.yggdrasil.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-30T16:06:40Z","receivedAt":"2005-04-30T16:06:40Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Sat, Apr 30, 2005 at 07:44:17AM -0700, Adam J. Richter wrote:\n>\n> \tI'd like to mention a couple of possible optimizations\n> for both the with and without compression approaches.\n> \n> \tIf you remove the gzip compression, then I imagine you could\n> do much of the IO of checking out files via sendfile, without\n> ever copying data to program space or even changing the program's\n> memory map.  There apparently exists a python sendfile module.\n> \n> \tIf this mercurial were written in C, much of the rest of\n> the IO could be optimized with mmap (to reduce copies) and writev\n> in the absense of a compression pass.  I don't know enough about\n> python to know if these optimizations are available.\n\nPython can do mmap, not sure about writev. \n\nBut I'm currently still in the \"keep it as simple as possible\" stage.\nThere's a bunch of room for optimization still, but if I do it all\nnow, it'll make things hard when I run into the next design change.\n\nAnd there's still some important core work that needs doing - checkout\nand commit need to be a subcase of the core merge code.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"}]}