{"thread":{"id":"317","subject":"Mercurial 0.3 vs git benchmarks","startedAt":"2005-04-26T00:41:11Z","lastAt":"2005-05-04T02:10:40Z","messageCount":116,"participants":["Matt Mackall","Daniel Phillips","Linus Torvalds","Mike Taht","Chris Wedgwood","Andreas Gal","Chris Mason","Magnus Damm","Bill Davidsen","H. Peter Anvin","Andrew Morton","Ingo Molnar","Florian Weimer","Thomas Glanzmann","Theodore Ts'o","Sean","Morten Welinder","Tom Lord","Noel Maddy","Andrea Arcangeli","Andrew Timberlake-Newell","Kevin Smith","Morgan Schweers","David Lang","Daniel Barkalow","Denys Duchier","Olivier Galibert","valdis.kletnieks@vt.edu","Daniel Jacobowitz","Ryan Anderson","Edgar Toernig","Sam Ravnborg","Kyle Moffett","David A. Wheeler"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1681","messageId":"20050426004111.GI21897@waste.org","threadId":"317","inReplyTo":null,"subject":"Mercurial 0.3 vs git benchmarks","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-26T00:41:11Z","receivedAt":"2005-04-26T00:41:11Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"This is to announce an updated version of Mercurial. Mercurial is a\nscalable, fast, distributed SCM that works in a model similar to BK\nand Monotone. It has functional clone/branch and pull/merge support\nand a working first pass implementation of network pull. It's also\nextremely small and hackable: it's about 1000 lines of code.\n\n http://selenic.com/mercurial/\n\nHere are the results of checking in the first 12 releases of Linux 2.6\ninto empty repositories for Mercurial v0.3 (hg) and git-pasky-0.7.\nThis is on my 512M Pentium M laptop. Times are in seconds.\n\n                 user         system       real        du -sh\nver    files   hg    git    hg    git    hg    git    hg   git\n\n2.6.0  15007 19.949 35.526 3.171 2.264 25.138 87.994 145M   89M\n2.6.1    998  5.906  4.018 0.573 0.464 10.267  5.937 146M   99M\n2.6.2   2370  9.696 13.051 0.752 0.652 12.970 15.167 150M  117M\n2.6.3   1906 10.528 11.509 0.816 0.639 18.406 14.318 152M  135M\n2.6.4   3185 11.140  7.380 0.997 0.731 15.265 12.412 156M  158M\n2.6.5   2261 10.961  6.939 0.843 0.640 20.564  8.522 158M  177M\n2.6.6   2642 11.803 10.043 0.870 0.678 22.360 11.515 162M  197M\n2.6.7   3772 18.411 15.243 1.189 0.915 32.397 21.498 165M  227M\n2.6.8   4604 20.922 16.054 1.406 1.041 39.622 25.056 172M  262M\n2.6.9   4712 19.306 12.145 1.421 1.102 35.663 24.958 179M  297M\n2.6.10  5384 23.022 18.154 1.393 1.182 40.947 32.085 186M  338M\n2.6.11  5662 27.211 19.138 1.791 1.253 42.605 31.902 193M  379M\n\ntar of .hg/   108175360\ntar of .git/  209385920\n\nFull-tree change status (no changes):\nhg:  real 0.799s  user 0.607s  sys 0.167s\ngit: real 0.124s  user 0.051s  sys 0.051s\n\nCheck-out time (2.6.0):\nhg:  real 34.084s  user 4.069s  sys 2.024s\ngit: real 30.487s  user 2.393s  sys 1.007s\n\nFull-tree working dir diff (2.6.0 base with 2.6.1 in working dir):\nhg:  real 4.920s  user 4.629s  sys 0.260s\ngit: real 3.531s  user 1.869s  sys 0.862s\n(this needed an update-cache --refresh on top of git commit, which\ntook another: real 2m52.764s  user 2.833s  sys 1.008s)\n\nMerge from 2.6.0 to 2.6.1:\nhg:  real 15.507s  user 6.175s  sys 0.442s\ngit: haven't quite figured this one out yet\n\nSome notes:\n\n- hg has a separate index file for each file checked in, which is why\n  the initial check-in is larger\n- this also means it touches twice as many files, typically\n- neither hg nor git quite fit in cache on my 512M laptop (nor does a\n  kernel compile), but the extra indexing makes hg's wall times a bit longer\n- hg does a form of delta compression, so each checkin requires\n  retrieving a previous version, checking its hash, doing a diff,\n  compressing it, and checking in the result\n- hg is written in pure Python\n\nDespite the above, it compares pretty well to git in speed and is\nquite a bit better in terms of storage space. By reducing the zlib\ncompression level, it could probably win across the board.\n\nThe size numbers will get dramatically more unbalanced with more\nhistory - a conversion of the history in BK to git is expected to take\nover 3G, which Mercurial may actually take less space due to storing\ncompressed binary forward-only deltas.\n\nWhile disk may be cheap, network bandwidth is not. Given that the\ncommon case usage of git will be to do network pulls, it will find\nmost of its speed wasted on waiting for the network. Mercurial will\nalmost certainly win here for typical developer usage as it can do\nefficient delta communication (though it currently doesn't attempt any\npipelining so suffers a bit in round trips).\n\nMore discussion about Mercurial's design can be found here:\n\n http://selenic.com/mercurial/notes.txt\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"1690","messageId":"200504252149.50735.phillips@istop.com","threadId":"317","inReplyTo":"20050426004111.GI21897@waste.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Daniel Phillips","fromEmail":"phillips@istop.com","sentAt":"2005-04-26T01:49:50Z","receivedAt":"2005-04-26T01:49:50Z","isPatch":false,"sender":{"key":"phillips@istop.com","avatar":null},"body":"On Monday 25 April 2005 20:41, Matt Mackall wrote:\n> Despite the above, it compares pretty well to git in speed and is\n> quite a bit better in terms of storage space. By reducing the zlib\n> compression level, it could probably win across the board.\n\nHi Matt,\n\nCongratulations on an impressive demo!  How about actually checking the \ncompression vs wall clock theory?  And I probably don't have to mention \npsyco...\n\nRegards,\n\nDaniel\n"},{"id":"1691","messageId":"Pine.LNX.4.58.0504251859550.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050426004111.GI21897@waste.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T02:08:28Z","receivedAt":"2005-04-26T02:08:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Apr 2005, Matt Mackall wrote:\n>\n> Here are the results of checking in the first 12 releases of Linux 2.6\n> into empty repositories for Mercurial v0.3 (hg) and git-pasky-0.7.\n> This is on my 512M Pentium M laptop. Times are in seconds.\n> \n>                  user         system       real        du -sh\n> ver    files   hg    git    hg    git    hg    git    hg   git\n> \n> 2.6.0  15007 19.949 35.526 3.171 2.264 25.138 87.994 145M   89M\n> 2.6.1    998  5.906  4.018 0.573 0.464 10.267  5.937 146M   99M\n> 2.6.2   2370  9.696 13.051 0.752 0.652 12.970 15.167 150M  117M\n> 2.6.3   1906 10.528 11.509 0.816 0.639 18.406 14.318 152M  135M\n> 2.6.4   3185 11.140  7.380 0.997 0.731 15.265 12.412 156M  158M\n> 2.6.5   2261 10.961  6.939 0.843 0.640 20.564  8.522 158M  177M\n> 2.6.6   2642 11.803 10.043 0.870 0.678 22.360 11.515 162M  197M\n> 2.6.7   3772 18.411 15.243 1.189 0.915 32.397 21.498 165M  227M\n> 2.6.8   4604 20.922 16.054 1.406 1.041 39.622 25.056 172M  262M\n> 2.6.9   4712 19.306 12.145 1.421 1.102 35.663 24.958 179M  297M\n> 2.6.10  5384 23.022 18.154 1.393 1.182 40.947 32.085 186M  338M\n> 2.6.11  5662 27.211 19.138 1.791 1.253 42.605 31.902 193M  379M\n\nThat time in checking things in is worrisome.\n\n\"git\" is basically linear in the size of the patch, which is what I want,\nsince most patches I work with are a couple of files at most. The patches\nyou are checking in are huge - I never actually work with a change that is\nas big as a whole release. I work with changes that are five files or\nsomething.\n\n\"hg\" seems to basically slow down the more patches you have applied. It's \nhard to tell from the limited test set, but look at \"user\" time. It seems \nto increase from 6 seconds to 27 seconds.\n\nTo make an interesting benchmark, try applying the first 200 patches in \nthe current git kernel archive. Can you do them three per second? THAT is \nthe thing you should optimize for, not checking in huge changes.\n\nIf you're checking in a change to 1000+ files, you're doing something\nwrong.\n\n> Full-tree working dir diff (2.6.0 base with 2.6.1 in working dir):\n> hg:  real 4.920s  user 4.629s  sys 0.260s\n> git: real 3.531s  user 1.869s  sys 0.862s\n> (this needed an update-cache --refresh on top of git commit, which\n> took another: real 2m52.764s  user 2.833s  sys 1.008s)\n\nYou're doing something wrong with git here. Why would you need to update \nyour cache?\n\n\t\t\tLinus\n"},{"id":"1694","messageId":"426DA7B5.2080204@timesys.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504251859550.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Mike Taht","fromEmail":"mike.taht@timesys.com","sentAt":"2005-04-26T02:30:13Z","receivedAt":"2005-04-26T02:30:13Z","isPatch":false,"sender":{"key":"mike.taht@timesys.com","avatar":null},"body":"Linus Torvalds wrote:\n > On Mon, 25 Apr 2005, Matt Mackall wrote:\n >\n >>Here are the results of checking in the first 12 releases of Linux 2.6\n >>into empty repositories for Mercurial v0.3 (hg) and git-pasky-0.7.\n >>This is on my 512M Pentium M laptop. Times are in seconds.\n\nOne difference is probably - mercurial appears to be using zlib's \n*default* compression of 6....\n\nusing zlib compression of 9 really impacts git...\n\nas per http://www.gelato.unsw.edu.au/archives/git/0504/1988.html\n\n >On a 700MHz p3, UDMA33, freebsd 5.3, ffs (soft updates) I get:\n\n >compressor | levels (size, time to compress, time to uncompress)\n >-----------+-------------------------------------------------------------------\n >gzip       | 9 (28M, 1:19, 30), 6 (28M, 31.7, 30), 3 (30M, 26.1,28.7)\n >           | 1 (31M, 23.6, 29.8)\n >bzip2      | 9 (27M, 2:14, 37.4) 6 (27M, 2:11, 38.8) 3 (27M, 2:10,38.3)\n >lzop       | 9 (32M, 2:15, 35.4) 7 (32M, 57.9, 40.3) 3 (39M, 36.0,44.4)\n\nas per setting GIT_COMPRESSION 3 rather than Z_BEST_COMPRESSION\n\nhttp://www.gelato.unsw.edu.au/archives/git/0504/1478.html\n\n\n>>\n>>                 user         system       real        du -sh\n>>ver    files   hg    git    hg    git    hg    git    hg   git\n>>\n>>2.6.0  15007 19.949 35.526 3.171 2.264 25.138 87.994 145M   89M\n>>2.6.1    998  5.906  4.018 0.573 0.464 10.267  5.937 146M   99M\n>>2.6.2   2370  9.696 13.051 0.752 0.652 12.970 15.167 150M  117M\n>>2.6.3   1906 10.528 11.509 0.816 0.639 18.406 14.318 152M  135M\n>>2.6.4   3185 11.140  7.380 0.997 0.731 15.265 12.412 156M  158M\n>>2.6.5   2261 10.961  6.939 0.843 0.640 20.564  8.522 158M  177M\n>>2.6.6   2642 11.803 10.043 0.870 0.678 22.360 11.515 162M  197M\n>>2.6.7   3772 18.411 15.243 1.189 0.915 32.397 21.498 165M  227M\n>>2.6.8   4604 20.922 16.054 1.406 1.041 39.622 25.056 172M  262M\n>>2.6.9   4712 19.306 12.145 1.421 1.102 35.663 24.958 179M  297M\n>>2.6.10  5384 23.022 18.154 1.393 1.182 40.947 32.085 186M  338M\n>>2.6.11  5662 27.211 19.138 1.791 1.253 42.605 31.902 193M  379M\n\n\n-- \n\nMike Taht\n\n\n   \"Imagination is more important than knowledge.\n\t-- Albert Einstein\"\n"},{"id":"1697","messageId":"Pine.LNX.4.58.0504251938210.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"426DA7B5.2080204@timesys.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T03:04:54Z","receivedAt":"2005-04-26T03:04:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Apr 2005, Mike Taht wrote:\n> \n> One difference is probably - mercurial appears to be using zlib's \n> *default* compression of 6....\n> \n> using zlib compression of 9 really impacts git...\n\nI agree that it will hurt for big changes, but since I really do believe \nthat most changes are just a couple of files, I don't believe it matters \nfor those. \n\nI forget what the exact numbers were, but I did some timings on plain\n\"gzip\", and it basically said that doing gzip on a medium-sized file was\nnot that different for -6 and -9. Why? Because most of the overhead was\nelsewhere ;)\n\nOh, well, I just re-created some numbers. This wasn't exactly what I did \nlast time I tested it, but it's conceptually the same thing:\n\n\ttorvalds@ppc970:~> time gzip -9 < v2.6/linux/kernel/sched.c > /dev/null \n\treal    0m0.018s\n\tuser    0m0.018s\n\tsys     0m0.000s\n\n\ttorvalds@ppc970:~> time gzip -6 < v2.6/linux/kernel/sched.c > /dev/null \n\treal    0m0.015s\n\tuser    0m0.013s\n\tsys     0m0.001s\n\nie there's a 0.003 second difference, which is certainly noticeable, and\nwould be hugely noticeable if you did a lot of these. But in my world-view\n(which is what git is optimized for), the common case is that you usually\nend up compressing maybe five-ten files, so the _compression_ overhead is\nnot that huge compared to all the other stuff.\n\nBut yes, testing git on big changes will test exactly the things that git\nisn't optimized for. I think git will normally hold up pretty well (ie it\nwill still beat anything that isn't designed for speed, and will be\ncomparable to things that _are_), but it's not what I'm interested in\noptimizing for.\n\nThat said - these days we can trivially change over to a \"zlib -6\" \ncompression, and nothing should ever notice. So if somebody wants to \ntest it, it should be fairly easy to just compare side-by-side: the \nresults should be identical.\n\nThe easiest test-case is Andrew's 198-patch patch-bomb on linux-kernel a \nfew weeks ago: they all apply cleanly to 2.6.12-rc2 (in order), and you \ncan use my \"dotest\" script to automate the test..\n\n\t\t\tLinus\n"},{"id":"1700","messageId":"Pine.LNX.4.58.0504252032500.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504251938210.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T04:00:05Z","receivedAt":"2005-04-26T04:00:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Apr 2005, Linus Torvalds wrote:\n> \n> The easiest test-case is Andrew's 198-patch patch-bomb on linux-kernel a \n> few weeks ago: they all apply cleanly to 2.6.12-rc2 (in order), and you \n> can use my \"dotest\" script to automate the test..\n\nOh, well. That was so trivial that I just did it:\n\nWith Z_BEST_COMPRESSION:\n\n\ttorvalds@ppc970:~/git-speed-1> ./script \n\tRemoving old tree\n\tCreating new tree\n\tInitializing db\n\tdefaulting to local storage area\n\tDoing sync\n\tInitial add\n\t\n\treal    0m37.526s\n\tuser    0m33.317s\n\tsys     0m3.816s\n\tInitial commit\n\tCommitting initial tree 0bba044c4ce775e45a88a51686b5d9f90697ea9d\n\t\n\treal    0m0.329s\n\tuser    0m0.152s\n\tsys     0m0.176s\n\tPatchbomb\n\t\n\treal    0m50.408s\n\tuser    0m18.933s\n\tsys     0m25.432s\n\nWith Z_DEFAULT_COMPRESSION:\n\n\ttorvalds@ppc970:~/git-speed-1> ./script \n\tRemoving old tree\n\tCreating new tree\n\tInitializing db\n\tdefaulting to local storage area\n\tDoing sync\n\tInitial add\n\t\n\treal    0m19.755s\n\tuser    0m15.719s\n\tsys     0m3.756s\n\tInitial commit\n\tCommitting initial tree 0bba044c4ce775e45a88a51686b5d9f90697ea9d\n\t\n\treal    0m0.337s\n\tuser    0m0.139s\n\tsys     0m0.197s\n\tPatchbomb\n\t\n\treal    0m50.465s\n\tuser    0m18.304s\n\tsys     0m25.567s\n\nie the \"initial add\" is almost twice as fast (because it spends most of\nthe time compressing _all_ the files), but the difference in applying 198\npatches is not noticeable at all (because the costs are all elsewhere).\n\nThat's 198 patches in less than a minute even with the highest\ncompression. That rocks.\n\nAnd don't try to make me explain why the patchbomb has any IO time at all,\nit should all have fit in the cache, but I think the writeback logic\nkicked in. Anyway, I tried it several times, and the real-time ends up \nfluctuating between 50-56 seconds, but the user/sys times are very stable, \nand end up being pretty much the same regardless of compression level.\n\nHere's the script, in case anybody cares:\n\n\t#!/bin/sh\n\techo Removing old tree\n\trm -rf linux-2.6.12-rc2\n\techo Creating new tree\n\tzcat < ~/v2.6/linux-2.6.12-rc2.tar.gz | tar xvf - > log\n\techo Initializing db\n\t( cd linux-2.6.12-rc2 ; init-db )\n\techo Doing sync\n\tsync\n\techo Initial add\n\ttime sh -c 'cd linux-2.6.12-rc2 && cat ../l | xargs update-cache --add --' >> log\n\techo Initial commit\n\ttime sh -c 'cd linux-2.6.12-rc2 && echo Initial commit | commit-tree \n\t$(write-tree) > .git/HEAD' >> log\n\techo Patchbomb\n\ttime sh -c 'cd linux-2.6.12-rc2 ; dotest ~/andrews-first-patchbomb' >> log\n\nand since the timing results were pretty much what I expected, I don't \nthink this changes _my_ opinion on anything. Yes, you can speed up commits \nwith Z_DEFAULT_COMPRESSION, but it's _not_ that big of a deal for my kind \nof model where you commit often, and commits are small.\n\nIt all boils down to:\n - huge commits are slowed down by compression overhead\n - I don't think huge commits really matter\n\nI mean, if it took 2 _hours_ to do the initial commit, I'd think it \nmatters. But when we're talking about less than a minute to create the \ninitial commit of a whole kernel archive, does it really make any \ndifference?\n\nAfter all, it's something you do _once_, and never again (unless you\nscript it to do performance testing ;)\n\nAnyway guys, feel free to test this on other machines. I bet there are\nlots of subtle performance differences between different filesystems and\nCPU architectures.. But the only hard numbers I have show that -9 isn't \nthat expensive.\n\n\t\t\tLinus\n"},{"id":"1701","messageId":"20050426040127.GK21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504251859550.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-26T04:01:27Z","receivedAt":"2005-04-26T04:01:27Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Mon, Apr 25, 2005 at 07:08:28PM -0700, Linus Torvalds wrote:\n> \n> \n> On Mon, 25 Apr 2005, Matt Mackall wrote:\n> >\n> > Here are the results of checking in the first 12 releases of Linux 2.6\n> > into empty repositories for Mercurial v0.3 (hg) and git-pasky-0.7.\n> > This is on my 512M Pentium M laptop. Times are in seconds.\n> > \n> >                  user         system       real        du -sh\n> > ver    files   hg    git    hg    git    hg    git    hg   git\n> > \n> > 2.6.0  15007 19.949 35.526 3.171 2.264 25.138 87.994 145M   89M\n> > 2.6.1    998  5.906  4.018 0.573 0.464 10.267  5.937 146M   99M\n> > 2.6.2   2370  9.696 13.051 0.752 0.652 12.970 15.167 150M  117M\n> > 2.6.3   1906 10.528 11.509 0.816 0.639 18.406 14.318 152M  135M\n> > 2.6.4   3185 11.140  7.380 0.997 0.731 15.265 12.412 156M  158M\n> > 2.6.5   2261 10.961  6.939 0.843 0.640 20.564  8.522 158M  177M\n> > 2.6.6   2642 11.803 10.043 0.870 0.678 22.360 11.515 162M  197M\n> > 2.6.7   3772 18.411 15.243 1.189 0.915 32.397 21.498 165M  227M\n> > 2.6.8   4604 20.922 16.054 1.406 1.041 39.622 25.056 172M  262M\n> > 2.6.9   4712 19.306 12.145 1.421 1.102 35.663 24.958 179M  297M\n> > 2.6.10  5384 23.022 18.154 1.393 1.182 40.947 32.085 186M  338M\n> > 2.6.11  5662 27.211 19.138 1.791 1.253 42.605 31.902 193M  379M\n> \n> That time in checking things in is worrisome.\n> \n> \"git\" is basically linear in the size of the patch, which is what I want,\n> since most patches I work with are a couple of files at most. The patches\n> you are checking in are huge - I never actually work with a change that is\n> as big as a whole release. I work with changes that are five files or\n> something.\n\nGit (and hg) commit time should be basically linear in the number of\nfiles touched, not the size of the patch.\n\n> \"hg\" seems to basically slow down the more patches you have applied. It's \n> hard to tell from the limited test set, but look at \"user\" time. It seems \n> to increase from 6 seconds to 27 seconds.\n\nAnd the number of files checked in grows from ~1000 to ~6000. Note\nthat git is growing from 4 to 19 seconds as well. Interestingly:\n\n19.138/4.018 = 4.76 (git time ratio)\n27.211/5.906 = 4.61 (hg time ratio)\n\nSo the scaling here is pretty similar.\n\n> To make an interesting benchmark, try applying the first 200 patches in \n> the current git kernel archive. Can you do them three per second? THAT is \n> the thing you should optimize for, not checking in huge changes.\n\nI'm not versant enough with git enough to know how but I'll give it a\nshot. Do you have the patches in an mbox, perchance? This is Andrew's\nx/198 patch bomb? It might be simpler for me to just apply everything\nin -mm to git and hg and compare times. Modulo python startup time, it\nshould be pretty similar.\n\nOh, and can you send me the script you used for your test with git?\n\n> If you're checking in a change to 1000+ files, you're doing something\n> wrong.\n\nThat's primarily to demonstrate the scalability and show the\ndivergence in repository sizes.\n\nHowever, it will not be uncommon for developers to pull/merge changes that\nlarge and the numbers here will be about the same for hg.\n\n> > Full-tree working dir diff (2.6.0 base with 2.6.1 in working dir):\n> > hg:  real 4.920s  user 4.629s  sys 0.260s\n> > git: real 3.531s  user 1.869s  sys 0.862s\n> > (this needed an update-cache --refresh on top of git commit, which\n> > took another: real 2m52.764s  user 2.833s  sys 1.008s)\n> \n> You're doing something wrong with git here. Why would you need to update \n> your cache?\n\nQuite possibly. Without it, I was getting a dump of a bunch of SHAs.\nI'm pretty git-ignorant, I've been focusing on something else for the\npast couple weeks.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"1702","messageId":"20050426040933.GA21178@taniwha.stupidest.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504251859550.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Chris Wedgwood","fromEmail":"cw@f00f.org","sentAt":"2005-04-26T04:09:33Z","receivedAt":"2005-04-26T04:09:33Z","isPatch":false,"sender":{"key":"cw@f00f.org","avatar":null},"body":"On Mon, Apr 25, 2005 at 07:08:28PM -0700, Linus Torvalds wrote:\n\n> If you're checking in a change to 1000+ files, you're doing\n> something wrong.\n\narch or subsystem merge?\n\n"},{"id":"1704","messageId":"Pine.LNX.4.58.0504252113210.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050426040127.GK21897@waste.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T04:20:12Z","receivedAt":"2005-04-26T04:20:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Apr 2005, Matt Mackall wrote:\n> \n> And the number of files checked in grows from ~1000 to ~6000. Note\n> that git is growing from 4 to 19 seconds as well.\n\nHeh,. I didn't much look at the git numbers, since I knew those were \nsupposed to be linear in the size of the patch...\n\n> I'm not versant enough with git enough to know how but I'll give it a\n> shot. Do you have the patches in an mbox, perchance? This is Andrew's\n> x/198 patch bomb?\n\nYes. I have my \"tools\" scripts for git in\n\n\tkernel.org:/pub/linux/kernel/people/torvalds/git-tools.git\n\nand I sent out the script I used to test the 2.6.12-rc2 + patches stuff in \nthe previous email, so you would just have to edit my mbox-applicator \ntools to work with hg and get comparable numbers.\n\n> It might be simpler for me to just apply everything\n> in -mm to git and hg and compare times.\n\nThat should work.\n\n> > You're doing something wrong with git here. Why would you need to update \n> > your cache?\n> \n> Quite possibly. Without it, I was getting a dump of a bunch of SHAs.\n> I'm pretty git-ignorant, I've been focusing on something else for the\n> past couple weeks.\n\nGetting a bunch of SHA's means that the file contents match, but that your \nindex file wasn't up-to-date, so git had to actually uncompress the object \nbacking store and _compare_ the file contents to notice.\n\nAnd I suspect that you may have done _all_ your numbers without ever\nhaving initialized the git index, in which case git will really suck raw\neggs, because git will basically always re-read every file (it will never\nrealize that they are up-to-date already).\n\nBasically, the theory of git operation is that the index file should\n_always_ be up-to-date.  Normally you don't have to do anything about it,\nsince the git helper tools will always just keep it that way, but if you\ndidn't, then..\n\n\t\tLinus\n"},{"id":"1706","messageId":"Pine.LNX.4.58.0504252117510.14838@sam.ics.uci.edu","threadId":"317","inReplyTo":"20050426040933.GA21178@taniwha.stupidest.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Andreas Gal","fromEmail":"gal@uci.edu","sentAt":"2005-04-26T04:22:10Z","receivedAt":"2005-04-26T04:22:10Z","isPatch":false,"sender":{"key":"gal@uci.edu","avatar":null},"body":"\nIf adding a new arch touches 1000+ files, you're doing something \n_very_ wrong. Plus, how often did that happen in the past 10 years? \n30 times?? Probably less.\n\nAndreas\n\nOn Mon, 25 Apr 2005, Chris Wedgwood wrote:\n\n> On Mon, Apr 25, 2005 at 07:08:28PM -0700, Linus Torvalds wrote:\n> \n> > If you're checking in a change to 1000+ files, you're doing\n> > something wrong.\n> \n> arch or subsystem merge?\n> \n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n> \n"},{"id":"1705","messageId":"Pine.LNX.4.58.0504252120211.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050426040933.GA21178@taniwha.stupidest.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T04:22:58Z","receivedAt":"2005-04-26T04:22:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Apr 2005, Chris Wedgwood wrote:\n>\n> On Mon, Apr 25, 2005 at 07:08:28PM -0700, Linus Torvalds wrote:\n> \n> > If you're checking in a change to 1000+ files, you're doing\n> > something wrong.\n> \n> arch or subsystem merge?\n\nNo, if it's a merge, you just suck in all the already-compressed objects.  \n\nYou never compress anything new - you get the objects, you update your\ntree index, and you're done. No overhead anywhere - a clean merge may\n_look_ like it's changing thousands of files, but it didn't change a\nsingle _object_ anywhere, it just re-arranged the objects and created a\nnew view of them.\n\nMost merges are literally just a tree-level thing. Sometimes you have to \ndo a content merge, but that tends to be a file or two.\n\n\t\t\tLinus\n"},{"id":"1729","messageId":"200504260713.26020.mason@suse.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504252032500.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-26T11:13:24Z","receivedAt":"2005-04-26T11:13:24Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Tuesday 26 April 2005 00:00, Linus Torvalds wrote:\n> On Mon, 25 Apr 2005, Linus Torvalds wrote:\n> > The easiest test-case is Andrew's 198-patch patch-bomb on linux-kernel a\n> > few weeks ago: they all apply cleanly to 2.6.12-rc2 (in order), and you\n> > can use my \"dotest\" script to automate the test..\n>\n> Oh, well. That was so trivial that I just did it:\n[ ... ]\n\n> ie the \"initial add\" is almost twice as fast (because it spends most of\n> the time compressing _all_ the files), but the difference in applying 198\n> patches is not noticeable at all (because the costs are all elsewhere).\n>\n> That's 198 patches in less than a minute even with the highest\n> compression. That rocks.\n\nThis agrees with my tests here, the time to apply patches is somewhat disk \nbound, even for the small 100 or 200 patch series.  The io should be coming \nfrom data=ordered, since the commits are still every 5 seconds or so.\n\n-chris\n"},{"id":"1740","messageId":"aec7e5c305042608095731d571@mail.gmail.com","threadId":"317","inReplyTo":"200504260713.26020.mason@suse.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Magnus Damm","fromEmail":"magnus.damm@gmail.com","sentAt":"2005-04-26T15:09:52Z","receivedAt":"2005-04-26T15:09:52Z","isPatch":false,"sender":{"key":"magnus.damm@gmail.com","avatar":null},"body":"On 4/26/05, Chris Mason <mason@suse.com> wrote:\n> This agrees with my tests here, the time to apply patches is somewhat disk\n> bound, even for the small 100 or 200 patch series.  The io should be coming\n> from data=ordered, since the commits are still every 5 seconds or so.\n\nYes, as long as you apply the patches to disk that is. I've hacked up\na small backend tool that applies patches to files kept in memory and\nuses a modifed rabin-karp search to match hunks. So you basically read\nonce and write once per file instead of moving data around for each\napplied patch. But it needs two passes.\n\nAnd no, the source code for the entire Linux kernel is not kept in\nmemory - you need a smart frontend to manage the file cache. Drop me a\nline if you are interested.\n\n/ magnus\n"},{"id":"1742","messageId":"200504261138.46339.mason@suse.com","threadId":"317","inReplyTo":"aec7e5c305042608095731d571@mail.gmail.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-26T15:38:45Z","receivedAt":"2005-04-26T15:38:45Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Tuesday 26 April 2005 11:09, Magnus Damm wrote:\n> On 4/26/05, Chris Mason <mason@suse.com> wrote:\n> > This agrees with my tests here, the time to apply patches is somewhat\n> > disk bound, even for the small 100 or 200 patch series.  The io should be\n> > coming from data=ordered, since the commits are still every 5 seconds or\n> > so.\n>\n> Yes, as long as you apply the patches to disk that is. I've hacked up\n> a small backend tool that applies patches to files kept in memory and\n> uses a modifed rabin-karp search to match hunks. So you basically read\n> once and write once per file instead of moving data around for each\n> applied patch. But it needs two passes.\n>\n> And no, the source code for the entire Linux kernel is not kept in\n> memory - you need a smart frontend to manage the file cache. Drop me a\n> line if you are interested.\n\nSorry, you've lost me.  Right now the cycle goes like this:\n\n1) patch reads patch file, reads source file, writes source file\n2) update-cache reads source file, writes git file\n\nWhich of those writes are you avoiding?  We have a smart way to manage the \ncache already for the source files...the vm does pretty well.  There's \nnothing to manage for the git files.  For the apply a bunch of patches \nworkload, they are write once, read never (except for the index).\n\n-chris\n\n"},{"id":"1745","messageId":"426E6845.2030008@tmr.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504251938210.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-04-26T16:11:49Z","receivedAt":"2005-04-26T16:11:49Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Mon, 25 Apr 2005, Mike Taht wrote:\n> \n>>One difference is probably - mercurial appears to be using zlib's \n>>*default* compression of 6....\n>>\n>>using zlib compression of 9 really impacts git...\n> \n> \n> I agree that it will hurt for big changes, but since I really do believe \n> that most changes are just a couple of files, I don't believe it matters \n> for those. \n> \n> I forget what the exact numbers were, but I did some timings on plain\n> \"gzip\", and it basically said that doing gzip on a medium-sized file was\n> not that different for -6 and -9. Why? Because most of the overhead was\n> elsewhere ;)\n\nCertainly not different in the overall numbers, but after trying gzip on \na bunch of various source files on 32 bit CPUs, from P-II to Xeon, it \nlooks as if after 7 the cpu jumps about 40% to 8, and another 30% to 9. \nNeither 8 nor 9 give any significant size improvement (< 2%).\n\nAgain, this is 32 bit CPU and just the gzip component, reading from \nstdin and writing to stdout which I hope gets directory operations out \nof the time measure.\n\nSample attached.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n\n\nComp level 1\n\nreal\t0m1.972s\nuser\t0m1.790s\nsys\t0m0.098s\n-rw-r--r--    1 davidsen  1792050 Apr 26 11:56 dummy.tar.gz\nComp level 2\n\nreal\t0m2.021s\nuser\t0m1.858s\nsys\t0m0.097s\n-rw-r--r--    1 davidsen  1737227 Apr 26 11:56 dummy.tar.gz\nComp level 3\n\nreal\t0m2.296s\nuser\t0m2.124s\nsys\t0m0.095s\n-rw-r--r--    1 davidsen  1697644 Apr 26 11:56 dummy.tar.gz\nComp level 4\n\nreal\t0m2.604s\nuser\t0m2.423s\nsys\t0m0.099s\n-rw-r--r--    1 davidsen  1593207 Apr 26 11:56 dummy.tar.gz\nComp level 5\n\nreal\t0m3.181s\nuser\t0m3.003s\nsys\t0m0.087s\n-rw-r--r--    1 davidsen  1549050 Apr 26 11:56 dummy.tar.gz\nComp level 6\n\nreal\t0m4.185s\nuser\t0m3.965s\nsys\t0m0.089s\n-rw-r--r--    1 davidsen  1531866 Apr 26 11:56 dummy.tar.gz\nComp level 7\n\nreal\t0m4.889s\nuser\t0m4.642s\nsys\t0m0.096s\n-rw-r--r--    1 davidsen  1524350 Apr 26 11:57 dummy.tar.gz\nComp level 8\n\nreal\t0m7.836s\nuser\t0m7.532s\nsys\t0m0.085s\n-rw-r--r--    1 davidsen  1513763 Apr 26 11:57 dummy.tar.gz\nComp level 9\n\nreal\t0m11.020s\nuser\t0m10.616s\nsys\t0m0.092s\n-rw-r--r--    1 davidsen  1511970 Apr 26 11:57 dummy.tar.gz\n"},{"id":"1746","messageId":"aec7e5c305042609231a5d3f0@mail.gmail.com","threadId":"317","inReplyTo":"200504261138.46339.mason@suse.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Magnus Damm","fromEmail":"magnus.damm@gmail.com","sentAt":"2005-04-26T16:23:11Z","receivedAt":"2005-04-26T16:23:11Z","isPatch":false,"sender":{"key":"magnus.damm@gmail.com","avatar":null},"body":"On 4/26/05, Chris Mason <mason@suse.com> wrote:\n> On Tuesday 26 April 2005 11:09, Magnus Damm wrote:\n> > On 4/26/05, Chris Mason <mason@suse.com> wrote:\n> > > This agrees with my tests here, the time to apply patches is somewhat\n> > > disk bound, even for the small 100 or 200 patch series.  The io should be\n> > > coming from data=ordered, since the commits are still every 5 seconds or\n> > > so.\n> >\n> > Yes, as long as you apply the patches to disk that is. I've hacked up\n> > a small backend tool that applies patches to files kept in memory and\n> > uses a modifed rabin-karp search to match hunks. So you basically read\n> > once and write once per file instead of moving data around for each\n> > applied patch. But it needs two passes.\n> >\n> > And no, the source code for the entire Linux kernel is not kept in\n> > memory - you need a smart frontend to manage the file cache. Drop me a\n> > line if you are interested.\n> \n> Sorry, you've lost me.  Right now the cycle goes like this:\n\nEhrm, maybe I'm way off. =)\n\n> 1) patch reads patch file, reads source file, writes source file\n> 2) update-cache reads source file, writes git file\n\nOk.\n\n> Which of those writes are you avoiding?  We have a smart way to manage the\n> cache already for the source files...the vm does pretty well.  There's\n> nothing to manage for the git files.  For the apply a bunch of patches\n> workload, they are write once, read never (except for the index).\n\nWell, maybe I misunderstood everything, but I thought you were\napplying a lot of patches and complained that it took a lot of time\ndue to the data order.\n\nWhen I applied a lot of patches to the kernel recently the cpu load\ndropped to zero after a while and the HD worked hard a sec or two and\nthen things came back again. My primitive guess is that it was because\nthe ext3 journal became full. To workaround this fact I started\nhacking on this in-memory patcher.\n\nIn the cycle above, I'm trying to speed up step 1:\nIf the patch modifies each source file multiple times (either using\nmultiple hunks or multiple ---/+++) then the lines below the hunk in\nthe source file will be moved multiple times. And if the source file\nis written to disk after each hunk or ---/+++ is applied then this\nwill generate a lot of writes that can be avoided if the entire patch\nprocedure is broken down into a first pass that analyzes the patches\nand a second pass that applies the patches and keeps source files in\nmemory.\n\nBut my rather trivial observation above is of course only suitable if\nyou have a lot of patches that should be applied and you are only\ninterested in the final version of the patched source files. If you\napply one patch at a time and import each source file as a new\nrevision then my little hack is probably not for you.\n\n/ magnus\n"},{"id":"1748","messageId":"Pine.LNX.4.58.0504260939440.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"200504260713.26020.mason@suse.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T16:42:33Z","receivedAt":"2005-04-26T16:42:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 26 Apr 2005, Chris Mason wrote:\n> \n> This agrees with my tests here, the time to apply patches is somewhat disk \n> bound, even for the small 100 or 200 patch series.  The io should be coming \n> from data=ordered, since the commits are still every 5 seconds or so.\n\nYes, ext3 really does suck in many ways.\n\nOne of my (least) favourite suckage is a process that does \"fsync\" on a\nsingle file (mail readers etc), which apparently causes ext3 to sync all\ndirty data, because it can only sync the whole log. So if you have stuff\nthat writes out things that aren't critical, it negatively affects\nsomething totally independent that _does_ care.\n\nI remember some early stuff showing that reiserfs was _much_ better for \nBK. I'd be willing to bet that's probably true for git too.\n\n\t\tLinus\n"},{"id":"1756","messageId":"200504261339.34680.mason@suse.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504260939440.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-26T17:39:33Z","receivedAt":"2005-04-26T17:39:33Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Tuesday 26 April 2005 12:42, Linus Torvalds wrote:\n> On Tue, 26 Apr 2005, Chris Mason wrote:\n> > This agrees with my tests here, the time to apply patches is somewhat\n> > disk bound, even for the small 100 or 200 patch series.  The io should be\n> > coming from data=ordered, since the commits are still every 5 seconds or\n> > so.\n>\n> Yes, ext3 really does suck in many ways.\n>\n> One of my (least) favourite suckage is a process that does \"fsync\" on a\n> single file (mail readers etc), which apparently causes ext3 to sync all\n> dirty data, because it can only sync the whole log. So if you have stuff\n> that writes out things that aren't critical, it negatively affects\n> something totally independent that _does_ care.\n>\n> I remember some early stuff showing that reiserfs was _much_ better for\n> BK. I'd be willing to bet that's probably true for git too.\n\nreiserfs shares the same basic data=ordered idea as ext3, so the fsync will do \nthe same on reiser as it does on ext3.  I do have code in there to try and \nkeep the data=ordered writeback a little less bursty than it is in ext3 so \nyou might not notice the fsync as much.\n\nI haven't compared reiser vs ext3 for git.  reiser tails should help \nperformance because once you read the object inode you've also got the data.  \nBut, I would expect the biggest help to come from mounting reiserfs -o \nalloc=skip_busy.  This basically allocates all new files one right after the \nother on disk regardless of which subdir they are in.  The effect is to time \norder most of your files.\n\nAs an example, here's the time to apply 300 patches on ext3.  This was with my \npacked patches applied, but vanilla git should show similar percentage \ndifferences.\n\ndata=writeback  32s\t\t\t\ndata=ordered    44s\n\nWith a long enough test, data=ordered should fall into the noise, but 10-40 \nsecond runs really show it.\n\n-chris\n"},{"id":"1762","messageId":"426E852A.40904@zytor.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504252032500.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-26T18:15:06Z","receivedAt":"2005-04-26T18:15:06Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> And don't try to make me explain why the patchbomb has any IO time at all,\n> it should all have fit in the cache, but I think the writeback logic\n> kicked in.\n\nThe default log size on ext3 is quite small.  Making the log larger \nprobably would have helped.\n\n\t-hpa\n"},{"id":"1764","messageId":"200504261418.18825.mason@suse.com","threadId":"317","inReplyTo":"aec7e5c305042609231a5d3f0@mail.gmail.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-26T18:18:18Z","receivedAt":"2005-04-26T18:18:18Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Tuesday 26 April 2005 12:23, Magnus Damm wrote:\n\n> Well, maybe I misunderstood everything, but I thought you were\n> applying a lot of patches and complained that it took a lot of time\n> due to the data order.\n>\n> When I applied a lot of patches to the kernel recently the cpu load\n> dropped to zero after a while and the HD worked hard a sec or two and\n> then things came back again. My primitive guess is that it was because\n> the ext3 journal became full. To workaround this fact I started\n> hacking on this in-memory patcher.\n\nIt looks like you'll only see the commits on ext3 when the log fills, and on \nreiser3 you'll see it every 5 seconds or when the log fills.  With the \ndefault mount options, both ext3 and reiser will flush the data blocks at the \nsame time they are writing the metadata.\n\nThe easiest way to get around this is to mount -o data=writeback on \next3/reiser, but you'll still have to wait for the data blocks eventually.  \n\n-chris\n"},{"id":"1780","messageId":"200504261552.24100.mason@suse.com","threadId":"317","inReplyTo":"200504261339.34680.mason@suse.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Chris Mason","fromEmail":"mason@suse.com","sentAt":"2005-04-26T19:52:23Z","receivedAt":"2005-04-26T19:52:23Z","isPatch":false,"sender":{"key":"mason@suse.com","avatar":null},"body":"On Tuesday 26 April 2005 13:39, Chris Mason wrote:\n\n> As an example, here's the time to apply 300 patches on ext3.  This was with\n> my packed patches applied, but vanilla git should show similar percentage\n> differences.\n>\n> data=writeback  32s\n> data=ordered    44s\n>\n> With a long enough test, data=ordered should fall into the noise, but 10-40\n> second runs really show it.\n\nI get much closer numbers if the patches directory is already in \ncache...data=ordered means more contention for the disk when trying to read \nthe patches.  \n\nIf the patches are hot in the cache data=writeback and data=ordered both take \nabout 30s.  You still see some writes in data=writeback, but these are mostly \nasync log commits.  \n\nThe same holds true for vanilla git as well, although it needs 1m7s to apply \nfrom a hot cache (sorry, couldn't resist the plug ;)\n\n-chris\n"},{"id":"1789","messageId":"426EA4D7.9080008@tmr.com","threadId":"317","inReplyTo":"426E852A.40904@zytor.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-04-26T20:30:15Z","receivedAt":"2005-04-26T20:30:15Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"H. Peter Anvin wrote:\n> Linus Torvalds wrote:\n> \n>>\n>> And don't try to make me explain why the patchbomb has any IO time at \n>> all,\n>> it should all have fit in the cache, but I think the writeback logic\n>> kicked in.\n> \n> \n> The default log size on ext3 is quite small.  Making the log larger \n> probably would have helped.\n\nExperience tells me that making the log larger does very good things for \nperformance in many load types. However, that fsync issue forcing the \nwrite of the whole log may get worse if there's a lot pending.\n\nI suspect this would be helped by noatime.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n"},{"id":"1798","messageId":"20050426135606.7b21a2e2.akpm@osdl.org","threadId":"317","inReplyTo":"aec7e5c305042609231a5d3f0@mail.gmail.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Andrew Morton","fromEmail":"akpm@osdl.org","sentAt":"2005-04-26T20:56:06Z","receivedAt":"2005-04-26T20:56:06Z","isPatch":false,"sender":{"key":"akpm@osdl.org","avatar":null},"body":"Magnus Damm <magnus.damm@gmail.com> wrote:\n>\n> My primitive guess is that it was because\n>  the ext3 journal became full.\n\nThe default ext3 journal size is inappropriately small, btw.  Normally you\nshould manually make it 128M or so, rather than 32M.  Unless you have a\nsmall amount of memory and/or a large number of filesystems, in which case\nthere might be problems with pinned memory.\n\nMounting as ext2 is a useful technique for determining whether the fs is\ngetting in the way.\n\n"},{"id":"1799","messageId":"Pine.LNX.4.58.0504261405050.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050426135606.7b21a2e2.akpm@osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-26T21:07:09Z","receivedAt":"2005-04-26T21:07:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 26 Apr 2005, Andrew Morton wrote:\n> \n> Mounting as ext2 is a useful technique for determining whether the fs is\n> getting in the way.\n\nWhat's the preferred way to try to convert a root filesystem to a bigger\njournal? Forcing \"rootfstype=ext2\" at boot and boot into single-user, and\nthen the appropriate magic tune2fs? Or what?\n\n\t\tLinus\n"},{"id":"1820","messageId":"426EC5A4.2090107@zytor.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504261405050.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-26T22:50:12Z","receivedAt":"2005-04-26T22:50:12Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Tue, 26 Apr 2005, Andrew Morton wrote:\n> \n>>Mounting as ext2 is a useful technique for determining whether the fs is\n>>getting in the way.\n> \n> \n> What's the preferred way to try to convert a root filesystem to a bigger\n> journal? Forcing \"rootfstype=ext2\" at boot and boot into single-user, and\n> then the appropriate magic tune2fs? Or what?\n> \n\nBoot single-user, \"remount -o ro,remount /\", \"tune2fs -J size=xxxM\" and \nreboot.\n\n\t-hpa\n"},{"id":"1821","messageId":"20050426155609.06e3ddcf.akpm@osdl.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504261405050.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Andrew Morton","fromEmail":"akpm@osdl.org","sentAt":"2005-04-26T22:56:09Z","receivedAt":"2005-04-26T22:56:09Z","isPatch":false,"sender":{"key":"akpm@osdl.org","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> wrote:\n>\n> \n> \n> On Tue, 26 Apr 2005, Andrew Morton wrote:\n> > \n> > Mounting as ext2 is a useful technique for determining whether the fs is\n> > getting in the way.\n> \n> What's the preferred way to try to convert a root filesystem to a bigger\n> journal? Forcing \"rootfstype=ext2\" at boot and boot into single-user, and\n> then the appropriate magic tune2fs? Or what?\n> \n\nGee, it's been ages.  umm,\n\n- umount the fs\n- tune2fs -O ^has_journal /dev/whatever\n- fsck -fy                              (to clean up the now-orphaned journal inode)\n- tune2fs -j -J size=nblocks    (normally 4k blocks)\n- mount the fs\n\n"},{"id":"1830","messageId":"426ED20B.9070706@zytor.com","threadId":"317","inReplyTo":"20050426155609.06e3ddcf.akpm@osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-26T23:43:07Z","receivedAt":"2005-04-26T23:43:07Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Andrew Morton wrote:\n> Linus Torvalds <torvalds@osdl.org> wrote:\n> \n>>\n>>\n>>On Tue, 26 Apr 2005, Andrew Morton wrote:\n>>\n>>>Mounting as ext2 is a useful technique for determining whether the fs is\n>>>getting in the way.\n>>\n>>What's the preferred way to try to convert a root filesystem to a bigger\n>>journal? Forcing \"rootfstype=ext2\" at boot and boot into single-user, and\n>>then the appropriate magic tune2fs? Or what?\n>>\n> \n> \n> Gee, it's been ages.  umm,\n> \n> - umount the fs\n> - tune2fs -O ^has_journal /dev/whatever\n> - fsck -fy                              (to clean up the now-orphaned journal inode)\n> - tune2fs -j -J size=nblocks    (normally 4k blocks)\n> - mount the fs\n> \n\nI think this is overkill, but should of course be safe.\n\nWhile you're doing this anyway, you might want to make sure you enable \n-O +dir_index and run fsck -D.\n\n\t-hpa\n"},{"id":"1855","messageId":"20050427063439.GA22014@elte.hu","threadId":"317","inReplyTo":"20050426135606.7b21a2e2.akpm@osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-04-27T06:34:39Z","receivedAt":"2005-04-27T06:34:39Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Andrew Morton <akpm@osdl.org> wrote:\n\n> Magnus Damm <magnus.damm@gmail.com> wrote:\n> >\n> > My primitive guess is that it was because\n> >  the ext3 journal became full.\n> \n> The default ext3 journal size is inappropriately small, btw.  Normally \n> you should manually make it 128M or so, rather than 32M.  Unless you \n> have a small amount of memory and/or a large number of filesystems, in \n> which case there might be problems with pinned memory.\n> \n> Mounting as ext2 is a useful technique for determining whether the fs \n> is getting in the way.\n\non ext3, when juggling patches and trees, the biggest performance boost \nfor me comes from adding noatime,nodiratime to the mount options in \n/etc/fstab:\n\n LABEL=/ / ext3 noatime,nodiratime,defaults 1 1\n\n\tIngo\n"},{"id":"1863","messageId":"871x8wb6w4.fsf@deneb.enyo.de","threadId":"317","inReplyTo":"426ED20B.9070706@zytor.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2005-04-27T15:01:47Z","receivedAt":"2005-04-27T15:01:47Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* H. Peter Anvin:\n\n> While you're doing this anyway, you might want to make sure you enable \n> -O +dir_index and run fsck -D.\n\nDirectory hashing has a negative impact on some applications (notably\ntar and unpatched mutt on large Maildir folders).  For git, it's a win\nbecause hashing destroys locality anyway.\n"},{"id":"1864","messageId":"20050427151357.GH1087@cip.informatik.uni-erlangen.de","threadId":"317","inReplyTo":"871x8wb6w4.fsf@deneb.enyo.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2005-04-27T15:13:57Z","receivedAt":"2005-04-27T15:13:57Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello,\n\n> Directory hashing has a negative impact on some applications (notably\n> tar and unpatched mutt on large Maildir folders).  For git, it's a win\n> because hashing destroys locality anyway.\n\nthis is inaccurate. Actually turning on directory hashing speeds-up big\nmaildirs a lot (tested with mutt-1.5.4 and higher with a maildir\ncontaining 30thousand messages). But in the mutt case you also have the\nheader cache[1] which speeds up a lot - with or without hashed\ndirectories. See also MEs comment[2] on this.\n\nFor tar I have no idea why it should slow down the operation, but maybe\nyou can enlighten us.\n\n\tThomas\n\n[1] http://wwwcip.informatik.uni-erlangen.de/~sithglan/mutt/\n\t- wait till TLR has released mutt-1.5.10\n\t- use mutt CVS HEAD\n\t- use mutt-1.5.9 + http://wwwcip.informatik.uni-erlangen.de/~sithglan/mutt/mutt-cvs-header-cache.29\n\t- and put the following in your .muttrc:\n\tset header_cache=/tmp/login-hcache\n\tset maildir_header_cache_verify=no\n\n[2] http://www.advogato.org/person/scandal/\n"},{"id":"1887","messageId":"426FDFCD.6000309@zytor.com","threadId":"317","inReplyTo":"20050427151357.GH1087@cip.informatik.uni-erlangen.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T18:54:05Z","receivedAt":"2005-04-27T18:54:05Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Thomas Glanzmann wrote:\n> \n> For tar I have no idea why it should slow down the operation, but maybe\n> you can enlighten us.\n> \n\nDirectory hashing slows down operations that do linear sweeps through \nthe filesystem reading every single file, simply because without \ndir_index, there is likely to be a correlation between inode order and \ndirectory order, whereas with dir_index, readdir() returns entries in \nhash order.\n\n\t-hpa\n"},{"id":"1888","messageId":"20050427190144.GA28848@cip.informatik.uni-erlangen.de","threadId":"317","inReplyTo":"426FDFCD.6000309@zytor.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2005-04-27T19:01:44Z","receivedAt":"2005-04-27T19:01:44Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello,\n\n> Directory hashing slows down operations that do linear sweeps through \n> the filesystem reading every single file, simply because without \n> dir_index, there is likely to be a correlation between inode order and \n> directory order, whereas with dir_index, readdir() returns entries in \n> hash order.\n\nthank you for the awareness training. Than mutt should be slower, too.\nMaybe I should repeat that tests.\n\n\tThomas\n"},{"id":"1894","messageId":"20050427195554.GA7793@thunk.org","threadId":"317","inReplyTo":"20050426155609.06e3ddcf.akpm@osdl.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2005-04-27T19:55:54Z","receivedAt":"2005-04-27T19:55:54Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Apr 26, 2005 at 03:56:09PM -0700, Andrew Morton wrote:\n> - umount the fs\n> - tune2fs -O ^has_journal /dev/whatever\n> - fsck -fy                              (to clean up the now-orphaned journal inode)\n\nUsing moderately recent versions of e2fsprogs, tune2fs will clean up\nthe journal inode, so there's no reason to do an fsck.  (Harmless, but\nit shouldn't be necessary and it takes time).\n\n> - tune2fs -j -J size=nblocks    (normally 4k blocks)\n\nThe argument to \"-J size\" is in megabytes, not in blocks.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"1895","messageId":"20050427195753.GB7793@thunk.org","threadId":"317","inReplyTo":"20050427190144.GA28848@cip.informatik.uni-erlangen.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2005-04-27T19:57:53Z","receivedAt":"2005-04-27T19:57:53Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Apr 27, 2005 at 09:01:44PM +0200, Thomas Glanzmann wrote:\n> Hello,\n> \n> > Directory hashing slows down operations that do linear sweeps through \n> > the filesystem reading every single file, simply because without \n> > dir_index, there is likely to be a correlation between inode order and \n> > directory order, whereas with dir_index, readdir() returns entries in \n> > hash order.\n> \n> thank you for the awareness training. Than mutt should be slower, too.\n> Maybe I should repeat that tests.\n\nIf you are using the mutt in Debian unstable, it has the patch applied\nwhich qsorts based on inode number returned from readdir(), which is\nwhy you may not have been seeing the problem.\n\nOr you can LD_PRELOAD the attached quick hack....\n\n\t\t\t\t\t\t- Ted\n\n/*\n * readdir accelerator\n *\n * (C) Copyright 2003, 2004 by Theodore Ts'o.\n *\n * Compile using the command:\n *\n * gcc -o spd_readdir.so -shared spd_readdir.c -ldl\n *\n * %Begin-Header%\n * This file may be redistributed under the terms of the GNU Public\n * License.\n * %End-Header%\n * \n */\n\n#define ALLOC_STEPSIZE\t100\n#define MAX_DIRSIZE\t0\n\n#define DEBUG\n\n#ifdef DEBUG\n#define DEBUG_DIR(x)\t{if (do_debug) { x; }}\n#else\n#define DEBUG_DIR(x)\n#endif\n\n#define _GNU_SOURCE\n#define __USE_LARGEFILE64\n\n#include <stdio.h>\n#include <unistd.h>\n#include <sys/types.h>\n#include <sys/stat.h>\n#include <stdlib.h>\n#include <string.h>\n#include <dirent.h>\n#include <errno.h>\n#include <dlfcn.h>\n\nstruct dirent_s {\n\tunsigned long long d_ino;\n\tlong long d_off;\n\tunsigned short int d_reclen;\n\tunsigned char d_type;\n\tchar *d_name;\n};\n\nstruct dir_s {\n\tDIR\t*dir;\n\tint\tnum;\n\tint\tmax;\n\tstruct dirent_s *dp;\n\tint\tpos;\n\tint\tfd;\n\tstruct dirent ret_dir;\n\tstruct dirent64 ret_dir64;\n};\n\nstatic int (*real_closedir)(DIR *dir) = 0;\nstatic DIR *(*real_opendir)(const char *name) = 0;\nstatic struct dirent *(*real_readdir)(DIR *dir) = 0;\nstatic struct dirent64 *(*real_readdir64)(DIR *dir) = 0;\nstatic off_t (*real_telldir)(DIR *dir) = 0;\nstatic void (*real_seekdir)(DIR *dir, off_t offset) = 0;\nstatic int (*real_dirfd)(DIR *dir) = 0;\nstatic unsigned long max_dirsize = MAX_DIRSIZE;\nstatic num_open = 0;\n#ifdef DEBUG\nstatic int do_debug = 0;\n#endif\n\nstatic void setup_ptr()\n{\n\tchar *cp;\n\n\treal_opendir = dlsym(RTLD_NEXT, \"opendir\");\n\treal_closedir = dlsym(RTLD_NEXT, \"closedir\");\n\treal_readdir = dlsym(RTLD_NEXT, \"readdir\");\n\treal_readdir64 = dlsym(RTLD_NEXT, \"readdir64\");\n\treal_telldir = dlsym(RTLD_NEXT, \"telldir\");\n\treal_seekdir = dlsym(RTLD_NEXT, \"seekdir\");\n\treal_dirfd = dlsym(RTLD_NEXT, \"dirfd\");\n\tif ((cp = getenv(\"SPD_READDIR_MAX_SIZE\")) != NULL) {\n\t\tmax_dirsize = atol(cp);\n\t}\n#ifdef DEBUG\n\tif (getenv(\"SPD_READDIR_DEBUG\"))\n\t\tdo_debug++;\n#endif\n}\n\nstatic void free_cached_dir(struct dir_s *dirstruct)\n{\n\tint i;\n\n\tif (!dirstruct->dp)\n\t\treturn;\n\n\tfor (i=0; i < dirstruct->num; i++) {\n\t\tfree(dirstruct->dp[i].d_name);\n\t}\n\tfree(dirstruct->dp);\n\tdirstruct->dp = 0;\n}\t\n\nstatic int ino_cmp(const void *a, const void *b)\n{\n\tconst struct dirent_s *ds_a = (const struct dirent_s *) a;\n\tconst struct dirent_s *ds_b = (const struct dirent_s *) b;\n\tino_t i_a, i_b;\n\t\n\ti_a = ds_a->d_ino;\n\ti_b = ds_b->d_ino;\n\n\tif (ds_a->d_name[0] == '.') {\n\t\tif (ds_a->d_name[1] == 0)\n\t\t\ti_a = 0;\n\t\telse if ((ds_a->d_name[1] == '.') && (ds_a->d_name[2] == 0))\n\t\t\ti_a = 1;\n\t}\n\tif (ds_b->d_name[0] == '.') {\n\t\tif (ds_b->d_name[1] == 0)\n\t\t\ti_b = 0;\n\t\telse if ((ds_b->d_name[1] == '.') && (ds_b->d_name[2] == 0))\n\t\t\ti_b = 1;\n\t}\n\n\treturn (i_a - i_b);\n}\n\n\nDIR *opendir(const char *name)\n{\n\tDIR *dir;\n\tstruct dir_s\t*dirstruct;\n\tstruct dirent_s *ds, *dnew;\n\tstruct dirent64 *d;\n\tstruct stat st;\n\n\tif (!real_opendir)\n\t\tsetup_ptr();\n\n\tDEBUG_DIR(printf(\"Opendir(%s) (%d open)\\n\", name, num_open++));\n\tdir = (*real_opendir)(name);\n\tif (!dir)\n\t\treturn NULL;\n\n\tdirstruct = malloc(sizeof(struct dir_s));\n\tif (!dirstruct) {\n\t\t(*real_closedir)(dir);\n\t\terrno = -ENOMEM;\n\t\treturn NULL;\n\t}\n\tdirstruct->num = 0;\n\tdirstruct->max = 0;\n\tdirstruct->dp = 0;\n\tdirstruct->pos = 0;\n\tdirstruct->dir = 0;\n\n\tif (max_dirsize && (stat(name, &st) == 0) && \n\t    (st.st_size > max_dirsize)) {\n\t\tDEBUG_DIR(printf(\"Directory size %ld, using direct readdir\\n\",\n\t\t\t\t st.st_size));\n\t\tdirstruct->dir = dir;\n\t\treturn (DIR *) dirstruct;\n\t}\n\n\twhile ((d = (*real_readdir64)(dir)) != NULL) {\n\t\tif (dirstruct->num >= dirstruct->max) {\n\t\t\tdirstruct->max += ALLOC_STEPSIZE;\n\t\t\tDEBUG_DIR(printf(\"Reallocating to size %d\\n\", \n\t\t\t\t\t dirstruct->max));\n\t\t\tdnew = realloc(dirstruct->dp, \n\t\t\t\t       dirstruct->max * sizeof(struct dir_s));\n\t\t\tif (!dnew)\n\t\t\t\tgoto nomem;\n\t\t\tdirstruct->dp = dnew;\n\t\t}\n\t\tds = &dirstruct->dp[dirstruct->num++];\n\t\tds->d_ino = d->d_ino;\n\t\tds->d_off = d->d_off;\n\t\tds->d_reclen = d->d_reclen;\n\t\tds->d_type = d->d_type;\n\t\tif ((ds->d_name = malloc(strlen(d->d_name)+1)) == NULL) {\n\t\t\tdirstruct->num--;\n\t\t\tgoto nomem;\n\t\t}\n\t\tstrcpy(ds->d_name, d->d_name);\n\t\tDEBUG_DIR(printf(\"readdir: %lu %s\\n\", \n\t\t\t\t (unsigned long) d->d_ino, d->d_name));\n\t}\n\tdirstruct->fd = dup((*real_dirfd)(dir));\n\t(*real_closedir)(dir);\n\tqsort(dirstruct->dp, dirstruct->num, sizeof(struct dirent_s), ino_cmp);\n\treturn ((DIR *) dirstruct);\nnomem:\n\tDEBUG_DIR(printf(\"No memory, backing off to direct readdir\\n\"));\n\tfree_cached_dir(dirstruct);\n\tdirstruct->dir = dir;\n\treturn ((DIR *) dirstruct);\n}\n\nint closedir(DIR *dir)\n{\n\tstruct dir_s\t*dirstruct = (struct dir_s *) dir;\n\n\tDEBUG_DIR(printf(\"Closedir (%d open)\\n\", --num_open));\n\tif (dirstruct->dir)\n\t\t(*real_closedir)(dirstruct->dir);\n\n\tif (dirstruct->fd >= 0)\n\t\tclose(dirstruct->fd);\n\tfree_cached_dir(dirstruct);\n\tfree(dirstruct);\n\treturn 0;\n}\n\nstruct dirent *readdir(DIR *dir)\n{\n\tstruct dir_s\t*dirstruct = (struct dir_s *) dir;\n\tstruct dirent_s *ds;\n\n\tif (dirstruct->dir)\n\t\treturn (*real_readdir)(dirstruct->dir);\n\n\tif (dirstruct->pos >= dirstruct->num)\n\t\treturn NULL;\n\n\tds = &dirstruct->dp[dirstruct->pos++];\n\tdirstruct->ret_dir.d_ino = ds->d_ino;\n\tdirstruct->ret_dir.d_off = ds->d_off;\n\tdirstruct->ret_dir.d_reclen = ds->d_reclen;\n\tdirstruct->ret_dir.d_type = ds->d_type;\n\tstrncpy(dirstruct->ret_dir.d_name, ds->d_name,\n\t\tsizeof(dirstruct->ret_dir.d_name));\n\n\treturn (&dirstruct->ret_dir);\n}\n\nstruct dirent64 *readdir64(DIR *dir)\n{\n\tstruct dir_s\t*dirstruct = (struct dir_s *) dir;\n\tstruct dirent_s *ds;\n\n\tif (dirstruct->dir)\n\t\treturn (*real_readdir64)(dirstruct->dir);\n\n\tif (dirstruct->pos >= dirstruct->num)\n\t\treturn NULL;\n\n\tds = &dirstruct->dp[dirstruct->pos++];\n\tdirstruct->ret_dir64.d_ino = ds->d_ino;\n\tdirstruct->ret_dir64.d_off = ds->d_off;\n\tdirstruct->ret_dir64.d_reclen = ds->d_reclen;\n\tdirstruct->ret_dir64.d_type = ds->d_type;\n\tstrncpy(dirstruct->ret_dir64.d_name, ds->d_name,\n\t\tsizeof(dirstruct->ret_dir64.d_name));\n\n\treturn (&dirstruct->ret_dir64);\n}\n\noff_t telldir(DIR *dir)\n{\n\tstruct dir_s\t*dirstruct = (struct dir_s *) dir;\n\n\tif (dirstruct->dir)\n\t\treturn (*real_telldir)(dirstruct->dir);\n\n\treturn ((off_t) dirstruct->pos);\n}\n\nvoid seekdir(DIR *dir, off_t offset)\n{\n\tstruct dir_s\t*dirstruct = (struct dir_s *) dir;\n\n\tif (dirstruct->dir) {\n\t\t(*real_seekdir)(dirstruct->dir, offset);\n\t\treturn;\n\t}\n\n\tdirstruct->pos = offset;\n}\n\nint dirfd(DIR *dir)\n{\n\tstruct dir_s\t*dirstruct = (struct dir_s *) dir;\n\n\tif (dirstruct->dir)\n\t\treturn (*real_dirfd)(dirstruct->dir);\n\n\treturn (dirstruct->fd);\n}\n"},{"id":"1896","messageId":"20050427200617.GI28848@cip.informatik.uni-erlangen.de","threadId":"317","inReplyTo":"20050427195753.GB7793@thunk.org","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2005-04-27T20:06:17Z","receivedAt":"2005-04-27T20:06:17Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello,\n\n> Or you can LD_PRELOAD the attached quick hack....\n\nnice one! I have to keep that around. :-)\n\nThanks,\n\tThomas\n"},{"id":"1899","messageId":"426FF799.4000501@zytor.com","threadId":"317","inReplyTo":"20050427190144.GA28848@cip.informatik.uni-erlangen.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T20:35:37Z","receivedAt":"2005-04-27T20:35:37Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Thomas Glanzmann wrote:\n> Hello,\n> \n> \n>>Directory hashing slows down operations that do linear sweeps through \n>>the filesystem reading every single file, simply because without \n>>dir_index, there is likely to be a correlation between inode order and \n>>directory order, whereas with dir_index, readdir() returns entries in \n>>hash order.\n> \n> \n> thank you for the awareness training. Than mutt should be slower, too.\n> Maybe I should repeat that tests.\n> \n\nOnly if you read every single file in each directory every time.  I \nthought mutt did header indexing and thus didn't need to do that.\n\n\t-hpa\n"},{"id":"1901","messageId":"20050427203917.GC12882@cip.informatik.uni-erlangen.de","threadId":"317","inReplyTo":"426FF799.4000501@zytor.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2005-04-27T20:39:18Z","receivedAt":"2005-04-27T20:39:18Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello,\n\n> Only if you read every single file in each directory every time.  I \n> thought mutt did header indexing and thus didn't need to do that.\n\nit does, but it is a very recent development (coming with the next\nrelease). Prior to this you need a patch, which has debian applied since\nsome time. And configure it. Otherwise *all* Maildir files we opened and\nparsed when a folder is entered.\n\n\tThomas\n"},{"id":"1902","messageId":"87br8054m3.fsf@deneb.enyo.de","threadId":"317","inReplyTo":"426FF799.4000501@zytor.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2005-04-27T20:47:32Z","receivedAt":"2005-04-27T20:47:32Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* H. Peter Anvin:\n\n> Only if you read every single file in each directory every time.  I \n> thought mutt did header indexing and thus didn't need to do that.\n\nThere was a patch for Mutt which implemented header indexing, but it\nwas buggy and had to be removed (from Debian).  After that, directory\nsorting (actually, it's a merge sort 8-) practically became mandatory\non ext3 with directory hashing.\n\nI think that in the meantime, the has been integrated into upstream\nCVS (I don't know if it's been released as a developer snapshot,\nthough).  The header indexing patch may have been revived for Debian,\nI think it was fixed recently.\n"},{"id":"1905","messageId":"874qds5489.fsf@deneb.enyo.de","threadId":"317","inReplyTo":"20050427190144.GA28848@cip.informatik.uni-erlangen.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2005-04-27T20:55:50Z","receivedAt":"2005-04-27T20:55:50Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* Thomas Glanzmann:\n\n>> Directory hashing slows down operations that do linear sweeps through \n>> the filesystem reading every single file, simply because without \n>> dir_index, there is likely to be a correlation between inode order and \n>> directory order, whereas with dir_index, readdir() returns entries in \n>> hash order.\n>\n> thank you for the awareness training. Than mutt should be slower, too.\n> Maybe I should repeat that tests.\n\nBenchmarks are actually a bit tricky because as far as I can tell,\nonce you hash the directories, they are tainted even if you mount your\nfile system with ext2.\n"},{"id":"1908","messageId":"426FFE58.4050901@zytor.com","threadId":"317","inReplyTo":"874qds5489.fsf@deneb.enyo.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-04-27T21:04:24Z","receivedAt":"2005-04-27T21:04:24Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Florian Weimer wrote:\n> \n> Benchmarks are actually a bit tricky because as far as I can tell,\n> once you hash the directories, they are tainted even if you mount your\n> file system with ext2.\n\nThat's what fsck -D is for.\n\n\t-hpa\n"},{"id":"1909","messageId":"87r7gw3p6p.fsf@deneb.enyo.de","threadId":"317","inReplyTo":"426FFE58.4050901@zytor.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2005-04-27T21:06:06Z","receivedAt":"2005-04-27T21:06:06Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* H. Peter Anvin:\n\n> Florian Weimer wrote:\n>> Benchmarks are actually a bit tricky because as far as I can tell,\n>> once you hash the directories, they are tainted even if you mount your\n>> file system with ext2.\n>\n> That's what fsck -D is for.\n\nAh, cool, I didn't know that it works the other way, too.  Thanks.\n"},{"id":"1910","messageId":"426FFFAB.1030005@tmr.com","threadId":"317","inReplyTo":"20050427063439.GA22014@elte.hu","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-04-27T21:10:03Z","receivedAt":"2005-04-27T21:10:03Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Ingo Molnar wrote:\n> * Andrew Morton <akpm@osdl.org> wrote:\n> \n> \n>>Magnus Damm <magnus.damm@gmail.com> wrote:\n>>\n>>>My primitive guess is that it was because\n>>> the ext3 journal became full.\n>>\n>>The default ext3 journal size is inappropriately small, btw.  Normally \n>>you should manually make it 128M or so, rather than 32M.  Unless you \n>>have a small amount of memory and/or a large number of filesystems, in \n>>which case there might be problems with pinned memory.\n>>\n>>Mounting as ext2 is a useful technique for determining whether the fs \n>>is getting in the way.\n> \n> \n> on ext3, when juggling patches and trees, the biggest performance boost \n> for me comes from adding noatime,nodiratime to the mount options in \n> /etc/fstab:\n> \n>  LABEL=/ / ext3 noatime,nodiratime,defaults 1 1\n\nI said much the same in another post, but noatime is not always what I \nreally want. How about a \"nojournalatime\" option, so the atime would be \nupdated at open and close, but not journaled at any other time. This \nwould reduce journal traffic but still allow an admin to tell if anyone \never uses a file. The info would be lost in a crash, but otherwise \npreserved just as it is for ext2. Might even be useful for ext2, not to \nwrite the atime, just track it in core.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n"},{"id":"1912","messageId":"20050427213250.GA8211@thunk.org","threadId":"317","inReplyTo":"87r7gw3p6p.fsf@deneb.enyo.de","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2005-04-27T21:32:50Z","receivedAt":"2005-04-27T21:32:50Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Apr 27, 2005 at 11:06:06PM +0200, Florian Weimer wrote:\n> * H. Peter Anvin:\n> \n> > Florian Weimer wrote:\n> >> Benchmarks are actually a bit tricky because as far as I can tell,\n> >> once you hash the directories, they are tainted even if you mount your\n> >> file system with ext2.\n> >\n> > That's what fsck -D is for.\n> \n> Ah, cool, I didn't know that it works the other way, too.  Thanks.\n\nIf htree support is disabled, e2fsck -D sorts by name, which was a\nsilly thing to do.  I should change it to sort by inode number instead\n(trivial patch).  This might not be a problem for the maildir format,\ngiven its naming convention.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"1914","messageId":"Pine.LNX.4.58.0504271431510.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"426FFFAB.1030005@tmr.com","subject":"Re: Mercurial 0.3 vs git benchmarks","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T21:39:03Z","receivedAt":"2005-04-27T21:39:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Bill Davidsen wrote:\n> \n> I said much the same in another post, but noatime is not always what I \n> really want.\n\n\"atime\" is really nasty for a filesystem. I don't know if anybody noticed, \nbut git already uses O_NOATIME to open all the object files, because if \nyou don't do that, then just looking at a full kernel tree (which has more \nthan a thousand subdirectories) will cause nasty IO patterns from just \nwriting back \"atime\" information for the \"tree\" objects we looked up.\n\nSo you can do (and git does) selective atime updates. It just requires a \nsmall amount of extra care. \n\n> How about a \"nojournalatime\" option, so the atime would be \n> updated at open and close, but not journaled at any other time.\n\nProbably a good idea. \n\n\t\tLinus\n"},{"id":"2118","messageId":"20050429060157.GS21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504251859550.18901@ppc970.osdl.org","subject":"Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T06:01:57Z","receivedAt":"2005-04-29T06:01:57Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Mon, Apr 25, 2005 at 07:08:28PM -0700, Linus Torvalds wrote:\n> \n> To make an interesting benchmark, try applying the first 200 patches in \n> the current git kernel archive. Can you do them three per second? THAT is \n> the thing you should optimize for, not checking in huge changes.\n\nOk, I've optimized for it a bit. This is basically:\n\n hg import -p1 -b ../broken-out `cat ../broken-out | grep -v #`\n\n ( latest code is at: http://selenic.com/mercurial/ )\n\nMy benchmark is to apply all 819 patches from -mm3 to 2.6.12-rc:\n\nhg:\n\nreal    3m22.075s\nuser    1m57.195s\nsys     0m14.068s\n\n819/(60+57.195 + 14.068) = 6.239 patches/second  user+sys\nrepository: before 167M after 173M (3.5% growth)\n\ngit:\n\nreal    2m58.568s\nuser    1m11.196s\nsys     0m50.144s\n\n819/(60+11.196+50.144) = 6.750 patches/second  user+sys\nrepository: before 102M after 154M (51% growth)\n\nAgain, pretty close, time-wise. My code is actually spending a fair\namount of time doing delta compression in Python, which accounts for\nmost of the extra user time. So I think I can optimize most of that\naway at some point. Interestingly hg is also using substantially less\nsystem time.\n\nWhat I'd like to highlight here is that git's repo is growing more\nthan 10 times faster. 52 megs is twice the size of a full kernel\ntarball. And that's going to be the bottleneck for network pull\noperations.\n\nThe fundamental problem I see with git is that the back-end has no\nconcept of the relation between files. This data is only present in\nchange nodes so you've got to potentially traverse all the commits to\nreconstruct a file's history. That's gonna be O(top-level changes)\nseeks. This introduces a number of problems:\n\n- no way to easily find previous revisions of a file\n  (being able to see when a particular change was introduced is a\n  pretty critical feature)\n- no way to do bandwidth-efficient delta transfer\n- no way to do efficient delta storage\n- no way to do merges based on the file's history[1]\n\nMercurial can grab look up and grab revisions of a file in O(1)\ntime/seeks. I haven't implemented annotate yet, but it can also be\ndone O(1) or O(file revisions).\n\n\n[1] This last one is interesting. If we've got a repository with files A\nand B:\n\nM   M1   M2\n\nAB\n |`-------v     M2 clones M\naB       AB     file A is change in mainline\n |`---v  AB'    file B is changed in M2\n |   aB / |     M1 clones M\n |   ab/  |     M1 changes B\n |   ab'  |     M1 merges from M2, changes to B conflict\n |    |  A'B'   M2 changes A\n  `---+--.|\n      |  a'B'   M2 merges from mainline, changes to A conflict\n      `--.|\n         ???    depending on which ancestor we choose, we will have\n\t        to redo A hand-merge, B hand-merge, or both\n                but if we look at the files independently, everything\n\t\tis fine\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2119","messageId":"3817.10.10.10.24.1114756831.squirrel@linux1","threadId":"317","inReplyTo":"20050429060157.GS21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2005-04-29T06:40:31Z","receivedAt":"2005-04-29T06:40:31Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Fri, April 29, 2005 2:01 am, Matt Mackall said:\n\n> What I'd like to highlight here is that git's repo is growing more\n> than 10 times faster. 52 megs is twice the size of a full kernel\n> tarball. And that's going to be the bottleneck for network pull\n> operations.\n\nThere isn't anything preventing optomized transfer protocols for git. \nGive it time.\n\n\n> The fundamental problem I see with git is that the back-end has no\n> concept of the relation between files. This data is only present in\n> change nodes so you've got to potentially traverse all the commits to\n> reconstruct a file's history. That's gonna be O(top-level changes)\n> seeks. This introduces a number of problems:\n>\n> - no way to easily find previous revisions of a file\n>   (being able to see when a particular change was introduced is a\n>   pretty critical feature)\n\nScanning back through the history is a linear operation and will quite\nlikely be just fine for many uses.   As others have pointed out, you can\ncache the result to improve subsequent lookups.\n\n\n> - no way to do bandwidth-efficient delta transfer\n\nThere's nothing preventing this in the longer term.  And you know, we're\nonly talking about a few megabytes per release.  We're not talking about\nvideo here.\n\n\n> - no way to do efficient delta storage\n\nThis has been discussed.  It is a recognized and accepted design\ntrade-off.  Disk is cheap.\n\n\n> - no way to do merges based on the file's history[1]\n\nWhat is preventing merges from looking back through the git history?\n\n\n\nThe fundamental design of git is essentially done, it is what it is.\n\nYour concearns are about performance rather than real limitations and it's\njust too damn early in the development process for that.  Frankly it's\namazing how good git is considering its age; it's already _way_ faster and\neasier to use than bk ever was for my use.\n\n\n> Mathematics is the supreme nostalgia of our time.\n\nI've been tying to figure out what this means for a while now <g>\n\n\nSean\n\n\n"},{"id":"2121","messageId":"20050429074043.GT21897@waste.org","threadId":"317","inReplyTo":"3817.10.10.10.24.1114756831.squirrel@linux1","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T07:40:43Z","receivedAt":"2005-04-29T07:40:43Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 02:40:31AM -0400, Sean wrote:\n> > - no way to do efficient delta storage\n> \n> This has been discussed.  It is a recognized and accepted design\n> trade-off.  Disk is cheap.\n\nThis trade-off FAILS, as my benchmarks against Mercurial have shown.\nIt trades 10x disk space for maybe 10% performance relative to my\napproach. Meanwhile, it makes a bunch of other things hard, namely the\nones I've listed. Yes, you can hack around them, but the back end will\nstill be bloated.\n\n> Your concearns are about performance rather than real limitations and it's\n> just too damn early in the development process for that.  Frankly it's\n> amazing how good git is considering its age; it's already _way_ faster and\n> easier to use than bk ever was for my use.\n\nMercurial is even younger (Linus had a few days' head start, not to\nmention a bunch of help), and it is already as fast as git, relatively\neasy to use, much simpler, and much more space and bandwidth\nefficient.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2122","messageId":"1680.10.10.10.24.1114764016.squirrel@linux1","threadId":"317","inReplyTo":"20050429074043.GT21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2005-04-29T08:40:16Z","receivedAt":"2005-04-29T08:40:16Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Fri, April 29, 2005 3:40 am, Matt Mackall said:\n\n> This trade-off FAILS, as my benchmarks against Mercurial have shown.\n> It trades 10x disk space for maybe 10% performance relative to my\n> approach. Meanwhile, it makes a bunch of other things hard, namely the\n> ones I've listed. Yes, you can hack around them, but the back end will\n> still be bloated.\n\nBut since performance can be seen as worth so much more than disk, this\nmight still be a good tradeoff, even given your numbers.\n\n\n> Mercurial is even younger (Linus had a few days' head start, not to\n> mention a bunch of help), and it is already as fast as git, relatively\n> easy to use, much simpler, and much more space and bandwidth\n> efficient.\n\n\nThere are some really nice things about the git design, not just\nperformance related.   However, i have a git repository going back to the\nstart of 2.4 and for my uses there aren't any performance problems.  (okay\nfsck-cache, gets oom killed but i suspect that can be fixed).\n\nNo _argument_ is going to change the fundamental design of git, it is what\nit is.  Git started out as just an interim fix and maybe that's all it\nwill turn out to be.  But it's working pretty well so far, with lots of\nroom for improvement over time, and in my estimation Linus has made a\npretty compelling argument for the design tradeoffs he's made.\n\nSean\n\n\n"},{"id":"2130","messageId":"Pine.LNX.4.58.0504290728090.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050429074043.GT21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-29T14:34:15Z","receivedAt":"2005-04-29T14:34:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 29 Apr 2005, Matt Mackall wrote:\n> \n> Mercurial is even younger (Linus had a few days' head start, not to\n> mention a bunch of help), and it is already as fast as git, relatively\n> easy to use, much simpler, and much more space and bandwidth\n> efficient.\n\nYou've not mentioned two out of my three design goals:\n - distribution\n - reliability/trustability\n\nie does mercurial do distributed merges, which git was designed for, and \ndoes mercurial notice single-bit errors in a reasonably secure manner, or \ncan people just mess with history willy-nilly?\n\nFor the latter, the cryptographic nature of sha1 is an added bonus - the\n_big_ issue is that it is a good hash, and an _exteremely_ effective CRC\nof the data. You can't mess up an archive and lie about it later. And if\nyou have random memory or filesystem corruption, it's not a \"shit happens\"  \nkind of situation - it's a \"uhhoh, we can catch it (and hopefully even fix\nit, thanks to distribution)\" thing.\n\nI had three design goals. \"disk space\" wasn't one of them, so you've\nconcentrated on only one so far in your arguments.\n\n\t\tLinus\n"},{"id":"2131","messageId":"118833cc05042908181d09bdfd@mail.gmail.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504290728090.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2005-04-29T15:18:20Z","receivedAt":"2005-04-29T15:18:20Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"> I had three design goals. \"disk space\" wasn't one of them\n\nAnd, if at some point it should become an issue, it's fixable.  Since\naccess to objects\nis fairly centralized and since they are immutable, it would be quite\nsimple to move\nan arbitrary selection of the objects into some other storage form\nwhich could take\nsimilarities between objects into account.\n\nIf you chose the selection of objects with care -- say those for files\nthat have changed\nmany times since -- it shouldn't even hurt performance of day-to-day\ntasks (which aren't\nlikely to ever need those objects).\n\nSo disk space and its cousin number-of-files are both when-and-if\nproblems.  And not\nscary ones at that.\n\nMorten\n"},{"id":"2133","messageId":"200504291544.IAA23584@emf.net","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504290728090.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T15:44:30Z","receivedAt":"2005-04-29T15:44:30Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n\n  > ie does mercurial do distributed merges, which git was designed for, and \n  > does mercurial notice single-bit errors in a reasonably secure manner, or \n  > can people just mess with history willy-nilly?\n\n  > For the latter, the cryptographic nature of sha1 is an added bonus - the\n  > _big_ issue is that it is a good hash, and an _exteremely_ effective CRC\n  > of the data. You can't mess up an archive and lie about it later.\n\nOn the other hand, you're asking people to sign whole trees and not just at\nfirst-import time but also for every change.\n\nThat's an impedence mismatch and undermines the security features of the\napproach you're taking and here is why:\n\nI shouldn't sign anything I haven't reviewed pretty carefully.  For\nthe kernel and in many other situations, it is too expensive to review\nthe whole tree.  Thus, the thing actually signed and the thing meant\nby the signature are not equal.  I sign a tree, in this system,\nbecause I think the right diffs and only the right diffs have been\napplied to it.   My signature is intended to mean, though, that I vouche\nfor the *diffs*, not the tree.\n\nIf I've changed five files, I should be signing a statement of:\n\n\t1) my belief about the identity of the immediate ancestor tree\n\t2) a robust summary of my changes, sufficient to recreate my\n\t   new tree given a faithful copy of the ancestor\n\nThat's a short enough amount of data that a human can really review it\nand thus it makes the signatures much more meaningful.\n\nProbably doesn't matter much other than in cases where a mainline\nis undergoing massive batch-patching based mostly on a web of trust.\n\nBut in that case --- someone or something generates purported diffs of\na tree; someone or something else scans those diffs and decides they\nlook good ---- and then on this basis, something distinct from\ndirectly using those diffs occurs.  The diffs were used to vette the\nchange; the signature asserts that a certain tree is a faithful result\nof applying those diffs.  Nothing checks that second assertion -- it's\ntaken on faith.\n\n-t\n\n"},{"id":"2136","messageId":"Pine.LNX.4.58.0504290854270.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"200504291544.IAA23584@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-29T15:58:37Z","receivedAt":"2005-04-29T15:58:37Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 29 Apr 2005, Tom Lord wrote:\n> \n> On the other hand, you're asking people to sign whole trees and not just at\n> first-import time but also for every change.\n\nI don't agree.\n\nSure, the commit determins the whole tree end result, but if you want to \nsign the _tree_, you can do so: just tag the actual _tree_ object as \"this \ntree has been verified to be bug-free and non-baby-seal-clubbing\".\n\nBut that's not what people do with tags. They sign a _commit_ object. And\nyes, the commit object points to the tree, but it also points to the whole\nhistory of other commit objects (and thus all historical trees etc), and \ntogether with just common sense it is very obvious that what you're really \nsigning is that \"point in time\".\n\nIf you want to clarify it, you can always just say so in the tag. Instead \nof saying \"I tag this as something I have verified every byte of\", you can \nsay \"this was what I released as xxx\", or \"this commit contains my change\" \nor something.\n\n> If I've changed five files, I should be signing a statement of:\n> \n> \t1) my belief about the identity of the immediate ancestor tree\n> \t2) a robust summary of my changes, sufficient to recreate my\n> \t   new tree given a faithful copy of the ancestor\n\nSo _do_ exactly that. You can say that in the tag you're signing.\n\n\t\t\tLinus\n"},{"id":"2141","messageId":"20050429163705.GU21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504290728090.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T16:37:05Z","receivedAt":"2005-04-29T16:37:05Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 07:34:15AM -0700, Linus Torvalds wrote:\n> \n> \n> On Fri, 29 Apr 2005, Matt Mackall wrote:\n> > \n> > Mercurial is even younger (Linus had a few days' head start, not to\n> > mention a bunch of help), and it is already as fast as git, relatively\n> > easy to use, much simpler, and much more space and bandwidth\n> > efficient.\n> \n> You've not mentioned two out of my three design goals:\n>  - distribution\n>  - reliability/trustability\n> \n> ie does mercurial do distributed merges, which git was designed for, and \n> does mercurial notice single-bit errors in a reasonably secure manner, or \n> can people just mess with history willy-nilly?\n\nDistribution: yes, it does BK/Monotone-style branching and merging.\nIn fact, these should be more \"correct\" than git as it has DAG\ninformation at the file level in the case where there are multiple\nancestors at the changeset graph level:\n\nM   M1   M2\n\nAB\n |`-------v     M2 clones M\naB       AB     file A is change in mainline\n |`---v  AB'    file B is changed in M2\n |   aB / |     M1 clones M\n |   ab/  |     M1 changes B\n |   ab'  |     M1 merges from M2, changes to B conflict\n |    |  A'B'   M2 changes A\n  `---+--.|\n      |  a'B'   M2 merges from mainline, changes to A conflict\n      `--.|\n         ???    depending on which ancestor we choose, we will have\n\t        to redo A hand-merge, B hand-merge, or both\n                but if we look at the files independently, everything\n\t\tis fine\n\n> For the latter, the cryptographic nature of sha1 is an added bonus - the\n> _big_ issue is that it is a good hash, and an _exteremely_ effective CRC\n> of the data. You can't mess up an archive and lie about it later. And if\n> you have random memory or filesystem corruption, it's not a \"shit happens\"  \n> kind of situation - it's a \"uhhoh, we can catch it (and hopefully even fix\n> it, thanks to distribution)\" thing.\n\nReliability/trustability: Mercurial is using a SHA1 hash as a checksum\nas well, much like Monotone and git. A changeset contains a hash of a\nmanifest which contains a hash of each file in the project, so you can\ndo things like sign the manifest hash (though I haven't implemented it\nyet. Making a backup is as simple as making a hardlink branch:\n\n mkdir backup\n cd backup\n hg branch ../linux  # takes about a second\n\n> I had three design goals. \"disk space\" wasn't one of them, so you've\n> concentrated on only one so far in your arguments.\n\nThat's because no one paid attention until I posted performance\nnumbers comparing it to git! Mercurial's goals are:\n\n- to scale to the kernel development process\n- to do clone/pull style development\n- to be efficient in CPU, memory, bandwidth, and disk space\n  for all the common SCM operations\n- to have strong repo integrity\n\nIt's been doing all that quite nicely since its first release.\nThe UI is also pretty straightforward:\n\nSetting up a Mercurial project:\n\n $ cd linux/\n $ hg init         # creates .hg\n $ hg status       # show changes between repo and working dir\n $ hg addremove    # add all unknown files and remove all missing files\n $ hg commit       # commit all changes, edit changelog entry\n\n Mercurial will look for a file named .hgignore in the root of your\n repository contains a set of regular expressions to ignore in file\n paths.\n\nMercurial commands:\n\n $ hg history          # show changesets\n $ hg log Makefile     # show commits per file\n $ hg diff             # generate a unidiff\n $ hg checkout         # check out the tip revision\n $ hg checkout <hash>  # check out a specified changeset\n $ hg add foo          # add a new file for the next commit\n $ hg remove bar       # mark a file as removed\n\nBranching and merging:\n\n $ cd ..\n $ mkdir linux-work\n $ cd linux-work\n $ hg branch ../linux        # create a new branch\n $ hg checkout               # populate the working directory\n $ <make changes>\n $ hg commit\n $ cd ../linux\n $ hg merge ../linux-work    # pull changesets from linux-work\n\nImporting patches:\n\n Fast:\n $ patch < ../p/foo.patch\n $ hg addremove\n $ hg commit\n\n Faster:\n $ patch < ../p/foo.patch\n $ hg commit `lsdiff -p1 ../p/foo.patch`\n\n Fastest:\n $ cat ../p/patchlist | xargs hg import -p1 -b ../p \n\nNetwork support:\n\n # export your .hg directory as a directory on your webserver\n foo$ ln -s .hg ~/public_html/hg-linux \n\n # merge changes from a remote machine\n bar$ hg init                              # create an empty repo\n bar$ hg merge http://foo/~user/hg-linux   # populate it\n bar$ <do some work>\n bar$ hg merge http://foo/~user/hg-linux   # resync\n\n This is just a proof of concept of grabbing byte ranges, and is not\n expected to perform well.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2142","messageId":"427264F2.1040609@tmr.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504290728090.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-04-29T16:46:42Z","receivedAt":"2005-04-29T16:46:42Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Fri, 29 Apr 2005, Matt Mackall wrote:\n> \n>>Mercurial is even younger (Linus had a few days' head start, not to\n>>mention a bunch of help), and it is already as fast as git, relatively\n>>easy to use, much simpler, and much more space and bandwidth\n>>efficient.\n> \n> \n> You've not mentioned two out of my three design goals:\n>  - distribution\n>  - reliability/trustability\n> \n> ie does mercurial do distributed merges, which git was designed for, and \n> does mercurial notice single-bit errors in a reasonably secure manner, or \n> can people just mess with history willy-nilly?\n> \n> For the latter, the cryptographic nature of sha1 is an added bonus - the\n> _big_ issue is that it is a good hash, and an _exteremely_ effective CRC\n> of the data. You can't mess up an archive and lie about it later. And if\n> you have random memory or filesystem corruption, it's not a \"shit happens\"  \n> kind of situation - it's a \"uhhoh, we can catch it (and hopefully even fix\n> it, thanks to distribution)\" thing.\n> \n> I had three design goals. \"disk space\" wasn't one of them, so you've\n> concentrated on only one so far in your arguments.\n\nReliability is a must have, but disk space matters in the real world if \nall other things are roughly equal. And bandwidth requirements are \ncertainly another real issue if they result in significant delay.\n\nIsn't the important thing  having the SCC reliable and easy to use, as \nin supports the things you want to do without jumping through hoops? One \nadvantage of Mercurial is that it can be the only major project for \nsomeone who seems to understand the problems, as opposed to taking the \ntime of someone (you) who has a load of other things in the fire. And if \nthere isn't time to do all the things you want, perhaps generating a \nwisj list and stepping back would be a good thing.\n\nIf you have the energy and time to stay with git, I'm sure it will be \ngreat, but you might want to provide input on Mercurial and let it run.\n\nPS: I don't think the performance difference is enough to constitute a \nreal advantage in either direction.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n"},{"id":"2143","messageId":"20050429165232.GV21897@waste.org","threadId":"317","inReplyTo":"118833cc05042908181d09bdfd@mail.gmail.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T16:52:32Z","receivedAt":"2005-04-29T16:52:32Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 11:18:20AM -0400, Morten Welinder wrote:\n> > I had three design goals. \"disk space\" wasn't one of them\n> \n> And, if at some point it should become an issue, it's fixable. Since\n> access to objects is fairly centralized and since they are\n> immutable, it would be quite simple to move an arbitrary selection\n> of the objects into some other storage form which could take\n> similarities between objects into account.\n\nThis is not a fix, this is a band-aid. A fix is fitting all the data\nin 10 times less space without sacrificing too much performance.\n\n> So disk space and its cousin number-of-files are both when-and-if\n> problems. And not scary ones at that.\n\nBut its sibling bandwidth _is_ a problem. The delta between 2.6.10 and\n2.6.11 in git terms will be much larger than a _full kernel tarball_.\nSimply checking in patch-2.6.11 on top of 2.6.10 as a single changeset\ntakes 41M. Break that into a thousand overlapping deltas (ie the way\nit is actually done) and it will be much larger.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2145","messageId":"Pine.LNX.4.58.0504291006450.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050429163705.GU21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-29T17:09:38Z","receivedAt":"2005-04-29T17:09:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 29 Apr 2005, Matt Mackall wrote:\n> \n> That's because no one paid attention until I posted performance\n> numbers comparing it to git! Mercurial's goals are:\n> \n> - to scale to the kernel development process\n> - to do clone/pull style development\n> - to be efficient in CPU, memory, bandwidth, and disk space\n>   for all the common SCM operations\n> - to have strong repo integrity\n\nOk, sounds good. Have you looked at how it scales over time, ie what \nhappens with files that have a lot of delta's?\n\nLet's see how these things work out..\n\n\t\tLinus\n"},{"id":"2146","messageId":"200504291734.KAA25263@emf.net","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504290854270.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T17:34:53Z","receivedAt":"2005-04-29T17:34:53Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n   From: Linus Torvalds <torvalds@osdl.org>\n\n   On Fri, 29 Apr 2005, Tom Lord wrote:\n   > \n   > On the other hand, you're asking people to sign whole trees and not\n   > just at first-import time but also for every change.\n\n   I don't agree.\n\nLet's be more precise, then.\n\n   Sure, the commit determins the whole tree end result, but if you want to \n   sign the _tree_, you can do so: just tag the actual _tree_ object as \"this \n   tree has been verified to be bug-free and non-baby-seal-clubbing\".\n\n   But that's not what people do with tags. They sign a _commit_ object. And\n   yes, the commit object points to the tree, but it also points to the whole\n   history of other commit objects (and thus all historical trees etc), and \n   together with just common sense it is very obvious that what you're really \n   signing is that \"point in time\".\n\n   If you want to clarify it, you can always just say so in the tag. Instead \n   of saying \"I tag this as something I have verified every byte of\", you can \n   say \"this was what I released as xxx\", or \"this commit contains my change\" \n   or something.\n\n\nA programmer publishing a change to the kernel has three pieces of \ndata in play.  They are:\nnn\n\t1) the ancestry of their modified tree\n\n\t2) the complete contents of their modified tree\n\n\t3) input data for a patching program (let's call it \"PATCH\")\n\t   which, at the very least, satisfies the equation:\n\n\t\tMOD_TREE = PATCH (this_diff, ORIG_TREE)\n\n\nAn upstream consumer, most often, is (should be) using (1) and (3).\nIn your system, the upstream consumer is given (1) and (2) and must\ncompute (3) for themselves.   The upstream consumer can also be provided\na signed version of (3) but if clients are mostly relying on the (2)\nthen there are multiple vulnerabilities there.\n\nThe set of pairs of type (1) and (2) is a dual space to the set of\npairs of type (1) and (3).   In that sense, it makes no mathematical\ndifference whether the programmer signs a {(1),(3)} pair or a {(1),(2)}\npair since either way, the other pair can be trivially derived.\n\nOn the other hand, signing documents which represent a {(1),(3)} pair\nwith robust accuracy is, in most cases, much much less expensive than\nsigning {(1),(2)} pairs with robust accuracy.  This is a bit like the\ndifference between \"*I* didn't set loose any mice in the house\" vs. \"I\nhave searched every corner of the house and swear there are no mice\nhere.\"\n\nAnother way to say that is that if someone gives me a signed {(1),(3)} pair\nI am likely to be much more confident that the signature represents\nan in-depth endorsement of the content as opposed to \"here is what the \ntools on my system happened to generate -- I hope it's what I meant\".\n\n\n   > If I've changed five files, I should be signing a statement of:\n   > \n   > \t1) my belief about the identity of the immediate ancestor tree\n   > \t2) a robust summary of my changes, sufficient to recreate my\n   > \t   new tree given a faithful copy of the ancestor\n\n   So _do_ exactly that. You can say that in the tag you're signing.\n\nWhich is pretty much exactly what Arch does except that, in the case\nof Arch, the signed diff is actually used directly to produce the mod\ntree.   There isn't a step in the process where a programmer reads\na purported diff, assumes that the signed tree accurately reflects\nthat diff, and then merges against the signed tree rather than the\ndiff.\n\nThe Arch approach also has the win that the signed diffs are useful\neven in the absense of the ORIG_TREE.  In the Arch world, we use this\nfor the form of selective merging that we call \"cherry-picking\" (picking\nthe desired changes from somebody else's line of development while \nskipping those which are not desired -- yet still being able to proceed\nwith bidirectional merging in a simple way).\n\nThe Arch approach also has the win that it amounts to delta-compression,\nsimplifying the design of efficient transports and helping to tame\nI/O bandwidth costs in the local case.   A lot of good features simply\n\"fall out\" of this approach.\n\nThis is kind of a yin-yang thing.  It's also valuable, in my view, to\nsign {(1),(2)} pairs.  It's also (in a more obscure way) valuable to\nsign {(1),(2),(3)} triples, especially if clients are regularly validating\nthem by making sure the triple describes a true diff application.\nSo one ultimately wants both functionalities, really.\n\nSigning trees which are really defined by a diff is a handy checksum\non the accuracy of tools which people are using but it isn't a substitute\nfor signing the diff itself and using it as the primary definition of \nthe tree it generates when combined with the immediate ancestor(s).\n\nSigning trees is also handy when that signature is then further linked\nto a set of binaries.\n\nSigning trees also is easier to implement and so gets git off the ground\nfaster.\n\nBut there's more to do, if a serious system is desired.\n\n-t\n\n"},{"id":"2149","messageId":"Pine.LNX.4.58.0504291051460.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"200504291734.KAA25263@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-29T17:56:30Z","receivedAt":"2005-04-29T17:56:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 29 Apr 2005, Tom Lord wrote:\n>\n> \t1) the ancestry of their modified tree\n> \n> \t2) the complete contents of their modified tree\n> \n> \t3) input data for a patching program (let's call it \"PATCH\")\n> \t   which, at the very least, satisfies the equation:\n> \n> \t\tMOD_TREE = PATCH (this_diff, ORIG_TREE)\n> \n> On the other hand, signing documents which represent a {(1),(3)} pair\n> with robust accuracy is, in most cases, much much less expensive than\n> signing {(1),(2)} pairs with robust accuracy. \n\nNot so.\n\nIt may be less expensive in your world, but that's the whole point of git: \nit's _not_ less expensive in the git world. \n\nIn the git world, 1 and 2 aren't even separate things. They go together. \nAnd you just sign it. End of story. It's so cheap to sign that it's not \neven funny.\n\nMore importantly, signing 3 is meaningless. 3 only makes sense with a \nknown starting point. You should never sign a patch without also saying \nwhat you're patching. \n\nAnd once you do that, 1+2 and 1+3 are _exactly_ the same thing.\n\nAnd since git always works on the 1+2 level, it would be inexcusably\nstupid to sign anything but that. 3 doesn't even exist per se, although \nit's obviously fully defined by 1+2.\n\nSo I don't see your point. You complain about git signing, but you \ncomplain on grounds that do not _exist_ in git, and then your alternative \n(1+3) which is senseless in a git world doesn't actually end up being \nanything really different - just more expensive.\n\n\t\tLinus\n"},{"id":"2150","messageId":"200504291808.LAA25870@emf.net","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504291051460.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T18:08:28Z","receivedAt":"2005-04-29T18:08:28Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n   From: Linus Torvalds <torvalds@osdl.org>\n\n   On Fri, 29 Apr 2005, Tom Lord wrote:\n   >\n   > \t1) the ancestry of their modified tree\n   > \n   > \t2) the complete contents of their modified tree\n   > \n   > \t3) input data for a patching program (let's call it \"PATCH\")\n   > \t   which, at the very least, satisfies the equation:\n   > \n   > \t\tMOD_TREE = PATCH (this_diff, ORIG_TREE)\n   > \n   > On the other hand, signing documents which represent a {(1),(3)} pair\n   > with robust accuracy is, in most cases, much much less expensive than\n   > signing {(1),(2)} pairs with robust accuracy. \n\n   Not so.\n\n   It may be less expensive in your world, but that's the whole point of git: \n   it's _not_ less expensive in the git world. \n\n   In the git world, 1 and 2 aren't even separate things. They go together. \n   And you just sign it. End of story. It's so cheap to sign that it's not \n   even funny.\n\nThe confusion here is that you are talking about computational complexity\nwhile I am talking about complexity measured in hours of labor.\n\nYou are assuming that the programmer generating the signature blindly \ntrusts the tool to generate the signed document accurately.   I am \nsaying that it should be tractable for human beings to read the documents\nthey are going to sign.\n\n\n   More importantly, signing 3 is meaningless. 3 only makes sense with a \n   known starting point. You should never sign a patch without also saying \n   what you're patching. \n\nI advocated signing a {(1),(3)} pair, not simply (3).\n\n   And once you do that, 1+2 and 1+3 are _exactly_ the same thing.\n\nI already spoke to that.\n\n   And since git always works on the 1+2 level, it would be inexcusably\n   stupid to sign anything but that. 3 doesn't even exist per se, although \n   it's obviously fully defined by 1+2.\n\n   So I don't see your point. You complain about git signing, but you \n   complain on grounds that do not _exist_ in git, and then your alternative \n   (1+3) which is senseless in a git world doesn't actually end up being \n   anything really different - just more expensive.\n\nI'm not sure what to suggest other than go back and read more carefully.\n\n-t\n\n"},{"id":"2154","messageId":"2712.10.10.10.24.1114799620.squirrel@linux1","threadId":"317","inReplyTo":"200504291808.LAA25870@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2005-04-29T18:33:40Z","receivedAt":"2005-04-29T18:33:40Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Fri, April 29, 2005 2:08 pm, Tom Lord said:\n\n> The confusion here is that you are talking about computational complexity\n> while I am talking about complexity measured in hours of labor.\n>\n> You are assuming that the programmer generating the signature blindly\n> trusts the tool to generate the signed document accurately.   I am\n> saying that it should be tractable for human beings to read the documents\n> they are going to sign.\n\n\nDevelopers obviously _do_ read the changes they submit to a project or\nthey would lose their trusted status.  That has absolutely nothing to do\nwith signing, it's the exact same way things work today, without sigs.\n\nIt's not \"blind trust\" to expect a script to reproducibly sign documents\nyou've decided to submit to a project.  The signature is not a QUALITY\nguarantee in and of itself.  It doesn't mean you have any additional\nresponsibility to remove all bugs before submitting.  Conversely, not\nsigning something doesn't mean you can submit crap.\n\nSee?  Signing something does not change the quality guarantee one way or\nthe other.  It does not put any additional demands on the developer, so\nit's fine to have an automated script do it.  It's just a way to avoid\nimpersonations.\n\nSean\n\n"},{"id":"2156","messageId":"200504291854.LAA26550@emf.net","threadId":"317","inReplyTo":"2712.10.10.10.24.1114799620.squirrel@linux1","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T18:54:09Z","receivedAt":"2005-04-29T18:54:09Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n   From: \"Sean\" <seanlkml@sympatico.ca>\n\n   On Fri, April 29, 2005 2:08 pm, Tom Lord said:\n\n   > The confusion here is that you are talking about computational complexity\n   > while I am talking about complexity measured in hours of labor.\n   >\n   > You are assuming that the programmer generating the signature blindly\n   > trusts the tool to generate the signed document accurately.   I am\n   > saying that it should be tractable for human beings to read the documents\n   > they are going to sign.\n\n\n   Developers obviously _do_ read the changes they submit to a project or\n   they would lose their trusted status.  That has absolutely nothing to do\n   with signing, it's the exact same way things work today, without sigs.\n\nNobody that I know is endorsing \"the way things work today\" as especially\nrobust.  Lots of people endorse it as successful in the marketplace and has\nhaving not failed horribly yet -- but that's not the same thing.\n\n\n   It's not \"blind trust\" to expect a script to reproducibly sign documents\n   you've decided to submit to a project.\n\nIt *is* blind trust to assume without further guarantees that the diff\nsomeone sends you (signed or not) describes a tree accurately unless\nthe tree in question is created by a local application of that diff.\n\nIn essense, `git' (today) wants *me* to trust that *you* have\ncorrectly applied that diff -- evidently in order to speed things up.\nIt makes remote users \"patch servers\", for no good reason.\n\nTriple signatures, signing both the name of the ancestor, the diff,\nand the resulting tree are the most robust because I can apply the\ndiff to the ancestor and then *verify* that it matches the signed\ntree.   But systems should neither ask users to sign something too large\nto read nor rely on signatures of things too large to read.\n\n\n   The signature is not a QUALITY\n   guarantee in and of itself.\n\nWhich has nothing to do with any of this except indirectly.\n\n   See?  Signing something does not change the quality guarantee one way or\n   the other.  It does not put any additional demands on the developer, so\n   it's fine to have an automated script do it.  It's just a way to avoid\n   impersonations.\n\nThe process should not rely on the security of every developer's\nmachine.  The process should not rely on simply trusting quality\ncontributors by reputation (e.g., most cons begin by establishing\ntrust and continue by relying inappropriately on\ntrust-without-verification).  This relates to why Linus'\nself-advertised process should be raising yellow and red cards all\nover the place: either he is wasting a huge amount of his own time and\nshould be largely replaced by an automated patch queue manager, or he\nis being trusted to do more than is humanly possible.\n\n-t\n"},{"id":"2157","messageId":"20050429191207.GX21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504291006450.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T19:12:08Z","receivedAt":"2005-04-29T19:12:08Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 10:09:38AM -0700, Linus Torvalds wrote:\n> \n> \n> On Fri, 29 Apr 2005, Matt Mackall wrote:\n> > \n> > That's because no one paid attention until I posted performance\n> > numbers comparing it to git! Mercurial's goals are:\n> > \n> > - to scale to the kernel development process\n> > - to do clone/pull style development\n> > - to be efficient in CPU, memory, bandwidth, and disk space\n> >   for all the common SCM operations\n> > - to have strong repo integrity\n> \n> Ok, sounds good. Have you looked at how it scales over time, ie what \n> happens with files that have a lot of delta's?\n\nI've done things like 10000 commits of a pair of revisions to printk.c\nand it maintains consistently high speed and compression throughout that\nrange. I've also done things like commit all 500 revisions of\nlinux/Makefile from bkcvs. This took a couple seconds and resulted in\nan 88k repo file (bkcvs takes 250k).\n\nI haven't tried the whole kernel history corpus yet, but I've\ncommitted all the 2.6 releases without any difficulties popping up and\nI've had handling >1M total file revisions in my head since I sat down\nto work on it. I'll maybe take a stab at a full history import next\nweek, if vacation doesn't interfere too much.\n\nOne downside Mercurial has is that long-lived repos can get fragmented on\ndisk. Things get defragmented to some extent as you go by doing COW on\nfiles that are shared between local branches clones. Also a complete\ndefrag is a simple cp -a or equivalent, so I think this is not a big\ndeal.\n\nHere's an excerpt from http://selenic.com/mercurial/notes.txt on how\nthe back-end works.\n\n---\n\nRevlogs:\n\nThe fundamental storage type in Mercurial is a \"revlog\". A revlog is\nthe set of all revisions to a file. Each revision is either stored\ncompressed in its entirety or as a compressed binary delta against the\nprevious version. The decision of when to store a full version is made\nbased on how much data would be needed to reconstruct the file. This\nlets us ensure that we never need to read huge amounts of data to\nreconstruct a file, regardless of how many revisions of it we store.\n\nIn fact, we should always be able to do it with a single read,\nprovided we know when and where to read. This is where the index comes\nin. Each revlog has an index containing a special hash (nodeid) of the\ntext, hashes for its parents, and where and how much of the revlog\ndata we need to read to reconstruct it. Thus, with one read of the\nindex and one read of the data, we can reconstruct any version in time\nproportional to the file size.\n\nSimilarly, revlogs and their indices are append-only. This means that\nadding a new version is also O(1) seeks.\n\nGenerally revlogs are used to represent revisions of files, but they\nalso are used to represent manifests and changesets.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2158","messageId":"2944.10.10.10.24.1114802002.squirrel@linux1","threadId":"317","inReplyTo":"200504291854.LAA26550@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2005-04-29T19:13:22Z","receivedAt":"2005-04-29T19:13:22Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Fri, April 29, 2005 2:54 pm, Tom Lord said:\n\n> The process should not rely on the security of every developer's\n> machine.  The process should not rely on simply trusting quality\n> contributors by reputation (e.g., most cons begin by establishing\n> trust and continue by relying inappropriately on\n> trust-without-verification).  This relates to why Linus'\n> self-advertised process should be raising yellow and red cards all\n> over the place: either he is wasting a huge amount of his own time and\n> should be largely replaced by an automated patch queue manager, or he\n> is being trusted to do more than is humanly possible.\n>\n\nAhh, you don't believe in the development model that has produced Linux! \nPersonally I do believe in it, so much so that I question the value of\nsignatures at the changeset level.  To me it doesn't matter where the code\ncame from just so long as it works.   Signatures are just a way to\nincrease the comfort level that the code has passed through a number of\npeople who have shown themselves to be relatively good auditors.  That's\nwhy I trust the code from my distribution of choice.  Everything is out in\nthe open anyway so it's much harder for a con man to do his thing.\n\nSean\n\n\n\n"},{"id":"2159","messageId":"200504291922.MAA27053@emf.net","threadId":"317","inReplyTo":"2944.10.10.10.24.1114802002.squirrel@linux1","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T19:22:19Z","receivedAt":"2005-04-29T19:22:19Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n\n  > Ahh, you don't believe in the development model that has produced Linux! \n  > Personally I do believe in it, so much so that I question the value of\n  > signatures at the changeset level.  To me it doesn't matter where the code\n  > came from just so long as it works.\n\nTo me, it doesn't matter where the code came from.  It's necessary\nbut not sufficient that it seems to work.  It's necessary that it's\nwell understood and has undergone only well understood changes.\n\nOn that last necessity, a *lot* of open source projects are quite\npathetic.  `git'-style use of signatures raises the bar, slightly, for\nwhere exploits can happen.  They also lower the bar for repudiation of\nbogus changes.\n\n   > Signatures are just a way to\n   > increase the comfort level that the code has passed through a number of\n   > people who have shown themselves to be relatively good auditors.  That's\n   > why I trust the code from my distribution of choice.  Everything is out in\n   > the open anyway so it's much harder for a con man to do his thing.\n\nOnly if the audience is proactively skeptical.\n\n-t\n\n"},{"id":"2160","messageId":"200504291928.MAA27145@emf.net","threadId":"317","inReplyTo":"2944.10.10.10.24.1114802002.squirrel@linux1","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T19:28:41Z","receivedAt":"2005-04-29T19:28:41Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\nThink of it this way:\n\n  (a) Joe, the mainline maintainer, gets a trusted message containing\n      a diff.\n\n  (b) Joe reads the diff, it makes great sense, he wants to merge.\n\n  (c) Joe downloads a tree.  Supposedly that tree is the result of\n      applying this diff.   The tree, not the diff, is used for\n      merging.\n\nYou can see the logical whole there... now the practical one:\n\n\n   (d) Joe is repeating (a..c) at an unfathomably high rate.\n       At a low rate, he could be double-checking enough that\n       that the diff-vs-tree problem isn't that serious.  But\n       at the rate he operates, exploits appear all along the\n       patch-flow pipeline because so much stuff goes unchecked.\n\n       Joe may be scan the changes he's merged before committing but,\n       if his rate is high, that scan *must*, out of biological and\n       physical necessity, be shallow.   Exploits can occur on the\n       submitter machine, in the communication channel, and on Joe's \n       machine.   Social exploits can occur because of the separation\n       between a submitter saying \"this is what I'm doing\" vs. the reality\n       of what the submitter is doing.\n\n-t\n\n"},{"id":"2163","messageId":"20050429194753.GA14222@uglybox.localnet","threadId":"317","inReplyTo":"200504291928.MAA27145@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Noel Maddy","fromEmail":"noel@zhtwn.com","sentAt":"2005-04-29T19:47:53Z","receivedAt":"2005-04-29T19:47:53Z","isPatch":false,"sender":{"key":"noel@zhtwn.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 12:28:41PM -0700, Tom Lord wrote:\n> \n> Think of it this way:\n> \n>   (a) Joe, the mainline maintainer, gets a trusted message containing\n>       a diff.\n> \n>   (b) Joe reads the diff, it makes great sense, he wants to merge.\n> \n>   (c) Joe downloads a tree.  Supposedly that tree is the result of\n>       applying this diff.   The tree, not the diff, is used for\n>       merging.\n\nCall me a naive git, but seems to me the \"git way\" is a little\ndifferent. It's tree-based rather than diff-based, and doesn't involve\npassing diffs around, right?\n\nThis is the process I'd expect:\n\n    (a)' Joe is notified of an update made to an external git tree\n\n    (b)' Joe pulls tree from the external git tree (signed by external\n         developer)\n\n    (c)' Joe reviews the (git-generated) diffs from his current\n\t     (trusted) tree to the new (signed) tree. If they pass\n         review, he merges the new versions into his tree, commits,\n         and signs his tree.\n\nThe logical hole that you point out is assuming that the diff is passed\nseparately from the tree rather than being directly generated from the\ncurrent maintainer tree and the signed remote tree.\n\nIf the diff is generated from the two signed trees, I don't see a hole.\n\nOr am I missing something?\n\n\n-- \nThe world's largest Internet database in the country.\n\t\t\t\t\t      -- Trading Times radio ad\n+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+\nNoel Maddy <noel@zhtwn.com>\n"},{"id":"2162","messageId":"Pine.LNX.4.58.0504291248210.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050429191207.GX21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-29T19:50:55Z","receivedAt":"2005-04-29T19:50:55Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 29 Apr 2005, Matt Mackall wrote:\n> \n> Here's an excerpt from http://selenic.com/mercurial/notes.txt on how\n> the back-end works.\n\nAny notes on how you maintain repository-level information?\n\nFor example, the expense in BK wasn't the single-file history, it was the\n_repository_ history, ie the \"ChangeSet\" file. Which grows quite slowly,\nbut because it _always_ grows, it ends up being quite big and expensive to\nparse after three years.\n\nIe do you have the git kind of \"independent trees/commits\", or do you \ncreate a revision history of those too?\n\n\t\tLinus\n"},{"id":"2166","messageId":"200504291954.MAA27561@emf.net","threadId":"317","inReplyTo":"20050429194753.GA14222@uglybox.localnet","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T19:54:19Z","receivedAt":"2005-04-29T19:54:19Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n\n  > Call me a naive git, but seems to me the \"git way\" is a little\n  > different. It's tree-based rather than diff-based, and doesn't involve\n  > passing diffs around, right?\n\nIsn't that a significant part of what I said?  Go back and read more\ncarefully, is my suggestion.\n\n  > Or am I missing something?\n\nVery much so.\n\n\n-t\n\n\n\n"},{"id":"2173","messageId":"000e01c54cf7$f61ee4a0$9b11a8c0@allianceoneinc.com","threadId":"317","inReplyTo":"200504291954.MAA27561@emf.net","subject":"RE: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrew Timberlake-Newell","fromEmail":"andrew.timberlake-newell@allianceoneinc.com","sentAt":"2005-04-29T20:13:50Z","receivedAt":"2005-04-29T20:13:50Z","isPatch":false,"sender":{"key":"andrew.timberlake-newell@allianceoneinc.com","avatar":null},"body":"Tom Lord responded to Noel Maddy: \n>   > Call me a naive git, but seems to me the \"git way\" is a little\n>   > different. It's tree-based rather than diff-based, and doesn't involve\n>   > passing diffs around, right?\n> \n> Isn't that a significant part of what I said?  Go back and read more\n> carefully, is my suggestion.\n\nIt looks to me like he did read carefully.\n\nThere were two different ideas:\n   TL)  Passing tree & diff and trusting diff to create tree\n   NM)  Passing tree and generating diff versus local tree for review\n\nMaybe I'm reading them wrong, but that certainly looks like what each was\nexpressing and they don't look like the same thing.\n\n\n"},{"id":"2178","messageId":"b8464fde050429131677ae06d1@mail.gmail.com","threadId":"317","inReplyTo":"200504291954.MAA27561@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Morgan Schweers","fromEmail":"mschweers@gmail.com","sentAt":"2005-04-29T20:16:44Z","receivedAt":"2005-04-29T20:16:44Z","isPatch":false,"sender":{"key":"mschweers@gmail.com","avatar":null},"body":"Greetings,\n\nOn 4/29/05, Tom Lord <lord@emf.net> wrote:\n> \n> \n>   > Call me a naive git, but seems to me the \"git way\" is a little\n>   > different. It's tree-based rather than diff-based, and doesn't involve\n>   > passing diffs around, right?\n> \n> Isn't that a significant part of what I said?  Go back and read more\n> carefully, is my suggestion.\n> \n>   > Or am I missing something?\n> \n> Very much so.\n\nIt doesn't appear that he is.  You appeared to predicate your argument\non the 'auditor' believing a diff looks good, but getting a tree\ninstead, that might not reflect the diff.\n\nInstead, in the git-world, the auditor actually gets a tree, and\nproduces the diff themselves, and then decides whether the diff looks\ngood enough to keep.\n\nThe argument about the high velocity of git-transfers causing the\ninability to check doesn't appear to apply here, because the\ndistributed development environment of Linux says that the\n'gatekeepers' ARE in fact validating the changes from people in their\narea of expertise are good (or are relying on sub-gatekeepers), and\nthen Linus is trusting them completely.\n\nThis seems like the methodology that has been used up until now via bk\npreviously.  Git doesn't change that, and in fact supports that method\nof development.\n\nYour further suggestion that Linus could be replaced by a\npatch-manager, in that case, got a chuckle from me at least, but the\nmore serious point is that Linus is necessary as the arbiter of who\nactually receives the absolute trust of a gatekeeper.  He is, in\neffect, a meta-gatekeeper.\n\n> -t\n\nIn reading this conversation, it seems you're looking for a more\nabsolute standard of trust than the kernel developers are working\nwith.  I believe this is an example of 'good enough' process being\naccepted, versus 'perfect' process.\n\n--  Morgan Schweers\n"},{"id":"2168","messageId":"20050429201957.GJ17379@opteron.random","threadId":"317","inReplyTo":"3817.10.10.10.24.1114756831.squirrel@linux1","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrea Arcangeli","fromEmail":"andrea@suse.de","sentAt":"2005-04-29T20:19:57Z","receivedAt":"2005-04-29T20:19:57Z","isPatch":false,"sender":{"key":"andrea@suse.de","avatar":null},"body":"On Fri, Apr 29, 2005 at 02:40:31AM -0400, Sean wrote:\n> There isn't anything preventing optomized transfer protocols for git. \n\nsuch a system might fall apart under load, converting on the fly from\ngit to network-optimized format sound quite expensive operation, even\nignorign the initial decompression of the payload. If something it\nshould be pre-converted to mercurial, so you checkout from mercurial and\nyou apply to local git.\n"},{"id":"2169","messageId":"20050429202117.GA15417@uglybox.localnet","threadId":"317","inReplyTo":"200504291954.MAA27561@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Noel Maddy","fromEmail":"noel@zhtwn.com","sentAt":"2005-04-29T20:21:17Z","receivedAt":"2005-04-29T20:21:17Z","isPatch":false,"sender":{"key":"noel@zhtwn.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 12:54:19PM -0700, Tom Lord wrote:\n> \n> \n>   > Call me a naive git, but seems to me the \"git way\" is a little\n>   > different. It's tree-based rather than diff-based, and doesn't involve\n>   > passing diffs around, right?\n> \n> Isn't that a significant part of what I said?  Go back and read more\n> carefully, is my suggestion.\n\nI'm trying to understand you. Please bear with me, and point out what\nI'm missing.\n\nYour example had Joe reviewing a signed diff, and then applying changes\nfrom a tree that \"supposedly\" had the diff applied correctly, but may\nhave been corrupted. If the tree was not an accurate representation of\napplying the diff, then the changes Joe applied to his tree will be\ndifferent than those that he reviewed.\n\nMy example had Joe downloading a remote signed tree, reviewing the changes\nlocally between his own trusted tree and the remote tree, and then\napplying them locally. Since the diffs are generated locally between the\ntwo trees, Joe is always reviewing the exact changes that will be\napplied to his tree.\n\nDoesn't this deal with the logical hole that you were pointing out in\nyour example? Or am I seeing a different \"logical hole\" than you are?\n\n\n-- \nA man who fears nothing is a man who loves nothing.  And if you love\nnothing, what joy is there in your life?\n\t\t\t\t\t -- King Arthur, \"First Knight\"\n+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+--+\nNoel Maddy <noel@zhtwn.com>\n"},{"id":"2170","messageId":"20050429202341.GB21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504291248210.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T20:23:41Z","receivedAt":"2005-04-29T20:23:41Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 12:50:55PM -0700, Linus Torvalds wrote:\n> \n> \n> On Fri, 29 Apr 2005, Matt Mackall wrote:\n> > \n> > Here's an excerpt from http://selenic.com/mercurial/notes.txt on how\n> > the back-end works.\n> \n> Any notes on how you maintain repository-level information?\n> \n> For example, the expense in BK wasn't the single-file history, it was the\n> _repository_ history, ie the \"ChangeSet\" file. Which grows quite slowly,\n> but because it _always_ grows, it ends up being quite big and expensive to\n> parse after three years.\n> \n> Ie do you have the git kind of \"independent trees/commits\", or do you \n> create a revision history of those too?\n\nThe changeset log (and everything else) has an external index. The\nindex is basically an array of (base, offset, length, parent1-hash,\nparent2-hash, my-hash). This has everything you need to reconstruct a\ngiven file revision with one seek/read into the data stream itself,\nand also everything you need for doing graph merging.\n\nThis is small enough (68 bytes, currently) that the index for a\nmillion changesets can be read into memory in a couple seconds or so,\neven in Python. It can also be mmapped and random accessed since the\nindex entries are fixed-sized. (And it's already stored big-endian.)\n\nSo you never have to read all the data. You also never need more than\na few indices in memory at once. And you never have to rewrite the\ndata (it's all append-only), except to do a bulk copy when you break a\nhardlink.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2175","messageId":"200504292026.NAA28131@emf.net","threadId":"317","inReplyTo":"000e01c54cf7$f61ee4a0$9b11a8c0@allianceoneinc.com","subject":"RE: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T20:26:31Z","receivedAt":"2005-04-29T20:26:31Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n\n  > It looks to me like he did read carefully.\n\n  > There were two different ideas:\n  >   TL)  Passing tree & diff and trusting diff to create tree\n  >   NM)  Passing tree and generating diff versus local tree for review\n\nWell, I guess *you* didn't read carefully.  I also spoke about the\nvalue of passing around triples: ancestry, diff, and tree.  The\nquestion is about linking signatures to things that humans can\nreasonably *intend* and be reasonably held accountable for, hence one\nof the values of signed diffs.  (I cited other practical reasons to\nvalue signed diffs and use them in specific ways, too.)\n\n-t\n"},{"id":"2177","messageId":"42729924.6090308@qualitycode.com","threadId":"317","inReplyTo":"200504291954.MAA27561@emf.net","subject":"Signed commit vulnerabilities? (was: Mercurial 0.4b vs git patchbomb benchmark)","fromName":"Kevin Smith","fromEmail":"yarcs@qualitycode.com","sentAt":"2005-04-29T20:29:24Z","receivedAt":"2005-04-29T20:29:24Z","isPatch":false,"sender":{"key":"yarcs@qualitycode.com","avatar":null},"body":"Tom Lord wrote:\n>   > Call me a naive git, but seems to me the \"git way\" is a little\n>   > different. It's tree-based rather than diff-based, and doesn't involve\n>   > passing diffs around, right?\n> \n> Isn't that a significant part of what I said?  Go back and read more\n> carefully, is my suggestion.\n> \n>   > Or am I missing something?\n> \n> Very much so.\n\nSo far, this is a frustrating conversation to watch. Here's my own\ninterpretation, presented to help the participants understand whether or\nnot their intended messages are getting through clearly.\n\nOriginally, Tom seemed to claim that the problem was that git requires\nyou to sign an entire tree, rather than a diff, even though the signer\nis only vouching for their diff.\n\nLinus responded by saying that a git signature of a tree would match\nthat description, but signing a commit is different. I think he claimed\nthat (by convention) signing a commit ONLY means you are signing the\nmost recent change, which turned tree A into tree B.\n\nTom then appeared to propose some specific attacks that could work\nagainst the git model. The precondition seems to be if the patch\nreceiver does not exhaustively analyze each and every patch. The\nreceiver trusts the contents based solely on who signed the commit object.\n\nOne category of attacks were that a computer or communication channel\nwas broken. It's not immediately clear to me how git's model contributes\nany weakness to these cases, compared to other signing strategies.\n\nThe other category of attack mentioned was social, such as a signer\ncreating a patch that claims to do one thing, but actually does another.\nAgain, I don't see how git is weaker in this case than any other tool.\n\nNoel then pointed out that in practice, someone receiving a signed\ncommit in git would view the commit comments and the diff, so the effect\nis similar to having the diff itself be signed.\n\nAnd that's where we are right now. So, from here, it looks like Tom\nneeds to be more specific about which attacks might be more effective\nagainst git's signing strategy than against signed diffs.\n\nKevin\n"},{"id":"2172","messageId":"20050429203027.GK17379@opteron.random","threadId":"317","inReplyTo":"20050429060157.GS21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrea Arcangeli","fromEmail":"andrea@suse.de","sentAt":"2005-04-29T20:30:27Z","receivedAt":"2005-04-29T20:30:27Z","isPatch":false,"sender":{"key":"andrea@suse.de","avatar":null},"body":"On Thu, Apr 28, 2005 at 11:01:57PM -0700, Matt Mackall wrote:\n> change nodes so you've got to potentially traverse all the commits to\n> reconstruct a file's history. That's gonna be O(top-level changes)\n> seeks. This introduces a number of problems:\n> \n> - no way to easily find previous revisions of a file\n>   (being able to see when a particular change was introduced is a\n>   pretty critical feature)\n> - no way to do bandwidth-efficient delta transfer\n> - no way to do efficient delta storage\n> - no way to do merges based on the file's history[1]\n\nAnd IMHO also no-way to implement a git-on-the-fly efficient network\nprotocol if tons of clients connects at the same time, it would be\ndosable etc... At the very least such a system would require an huge\namount of ram. So I see the only efficient way to design a network\nprotocol for git not to use git, but to import the data into mercurial\nand to implement the network protocol on top of mercurial.\n\nThe one downside is that git is sort of rock solid in the way it stores\ndata on disk, it makes rsync usage trivial too, the git fsck is reliable\nand you can just sign the hash of the root of the tree and you sign\neverything including file contents. And of course the checkin is\nabsolutely trivial and fast too.\n\nWith a more efficient diff-based storage like mercurial we'd be losing\nthose fsck properties etc.. but those reliability properties don't worth\nthe network and disk space they take IMHO, and the checkin time\nshouldn't be substantially different (still running in O(1) when\nappending at the head). And we could always store the hash of the\nchangeset, to give it some basic self-checking.\n\nI give extreme value in a SCM in how efficiently it can represent the\nwhole tree for both network downloads and backups too. Being able to\nstore the whole history of 2.5 in < 100M is a very valuable feature\nIMHO, much more valuable than to be able to sign the root.\n\nAlso don't get me wrong, I'm _very_ happy about git too, but I just\nhappen to prefer mercurial storage (I would never use git for anything\nbut the kernel, just like I wasn't using arch for similar reasons).\n"},{"id":"2181","messageId":"20050429203959.GC21897@waste.org","threadId":"317","inReplyTo":"20050429203027.GK17379@opteron.random","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T20:39:59Z","receivedAt":"2005-04-29T20:39:59Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 10:30:27PM +0200, Andrea Arcangeli wrote:\n> On Thu, Apr 28, 2005 at 11:01:57PM -0700, Matt Mackall wrote:\n> > change nodes so you've got to potentially traverse all the commits to\n> > reconstruct a file's history. That's gonna be O(top-level changes)\n> > seeks. This introduces a number of problems:\n> > \n> > - no way to easily find previous revisions of a file\n> >   (being able to see when a particular change was introduced is a\n> >   pretty critical feature)\n> > - no way to do bandwidth-efficient delta transfer\n> > - no way to do efficient delta storage\n> > - no way to do merges based on the file's history[1]\n> \n> And IMHO also no-way to implement a git-on-the-fly efficient network\n> protocol if tons of clients connects at the same time, it would be\n> dosable etc... At the very least such a system would require an huge\n> amount of ram. So I see the only efficient way to design a network\n> protocol for git not to use git, but to import the data into mercurial\n> and to implement the network protocol on top of mercurial.\n> \n> The one downside is that git is sort of rock solid in the way it stores\n> data on disk, it makes rsync usage trivial too, the git fsck is reliable\n> and you can just sign the hash of the root of the tree and you sign\n> everything including file contents. And of course the checkin is\n> absolutely trivial and fast too.\n\nMercurial is ammenable to rsync provided you devote a read-only\nrepository to it on the client side. In other words, you rsync from\nkernel.org/mercurial/linus to local/linus and then you merge from\nlocal/linus to your own branch. Mercurial's hashing hierarchy is\nsimilar to git's (and Monotone's), so you can sign a single hash of\nthe tree as well.\n\n> With a more efficient diff-based storage like mercurial we'd be losing\n> those fsck properties etc.. but those reliability properties don't worth\n> the network and disk space they take IMHO, and the checkin time\n> shouldn't be substantially different (still running in O(1) when\n> appending at the head). And we could always store the hash of the\n> changeset, to give it some basic self-checking.\n\nI think I can implement a decent repository check similar to git, it's\njust not been a priority.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2182","messageId":"Pine.LNX.4.62.0504291333550.7439@qynat.qvtvafvgr.pbz","threadId":"317","inReplyTo":"20050429202117.GA15417@uglybox.localnet","subject":"git network protocol","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2005-04-29T20:42:17Z","receivedAt":"2005-04-29T20:42:17Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"would it make sense for the network git protocol to be something along the \nlines of\n\nclient contacts server and sends\nthe tag you want to sync with (defaults to head)\nthe local index file\n\nthen the server can use the git tools locally to figure out what objects \nneed to be sent to do the merge and only send those objects.\n\nno this isn't as efficiant as only sending diffs, but it avoids sending \nany objects that aren't needed (which would be sent if you just did a \nstraight rsync)\n\nDavid Lang\n\n-- \nThere are two ways of constructing a software design. One way is to make it so simple that there are obviously no deficiencies. And the other way is to make it so complicated that there are no obvious deficiencies.\n  -- C.A.R. Hoare\n"},{"id":"2183","messageId":"200504292044.NAA28429@emf.net","threadId":"317","inReplyTo":"20050429202117.GA15417@uglybox.localnet","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Tom Lord","fromEmail":"lord@emf.net","sentAt":"2005-04-29T20:44:03Z","receivedAt":"2005-04-29T20:44:03Z","isPatch":false,"sender":{"key":"lord@emf.net","avatar":null},"body":"\n\n  > Your example had Joe reviewing a signed diff, and then applying changes\n  > from a tree that \"supposedly\" had the diff applied correctly, but may\n  > have been corrupted. If the tree was not an accurate representation of\n  > applying the diff, then the changes Joe applied to his tree will be\n  > different than those that he reviewed.\n\nThat's right.   I'm saying that Joe needn't rely on the tree at all since\nhe should be having his tools verify its contents anyway.  Given that, \nhe may as well have his tools *generate* the tree.  Having generated the tree,\nit's gravy to then verify that it matches the tree the submitter thought he\nwas sending -- that's a *secondary* checksum where `git' currently uses\nit as primary.\n\n\n  > My example had Joe downloading a remote signed tree, reviewing the changes\n  > locally between his own trusted tree and the remote tree, \n\nIn the real world, that \"review\" step is the weak link.  When it goes\nwrong, the first step is to make sure we are reviewing a tree everyone\ninvolved *intended* -- and it's only with signed diffs adding up to\nthat tree that we get there.\n\n-t\n"},{"id":"2185","messageId":"Pine.LNX.4.58.0504291338540.18901@ppc970.osdl.org","threadId":"317","inReplyTo":"20050429202341.GB21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-29T20:49:18Z","receivedAt":"2005-04-29T20:49:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 29 Apr 2005, Matt Mackall wrote:\n> \n> The changeset log (and everything else) has an external index.\n\nI don't actually know exactly how the BK changeset file works, but your \nexplanation really sounds _very_ much like it.\n\nI didn't want to do anything that even smelled of BK. Of course, part of\nmy reason for that is that I didn't feel comfortable with a delta model at\nall (I wouldn't know where to start, and I hate how they always end up\nhaving different rules for \"delta\"ble and \"non-delta\"ble objects).\n\nBut another was that exactly since I've been using BK for so long, I\nwanted to make sure that my model just emulated the way I've been _using_\nBK, rather than any BK technical details.\n\nSo it sounds like it could work fine, but it in fact sounds so much like \nthe ChangeSet file that I'd personally not have done it that way. \n\n\t\t\tLinus\n"},{"id":"2198","messageId":"001401c54cfe$061375f0$9b11a8c0@allianceoneinc.com","threadId":"317","inReplyTo":"200504292026.NAA28131@emf.net","subject":"RE: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrew Timberlake-Newell","fromEmail":"andrew.timberlake-newell@allianceoneinc.com","sentAt":"2005-04-29T20:57:15Z","receivedAt":"2005-04-29T20:57:15Z","isPatch":false,"sender":{"key":"andrew.timberlake-newell@allianceoneinc.com","avatar":null},"body":">   > It looks to me like he did read carefully.\n> \n>   > There were two different ideas:\n>   >   TL)  Passing tree & diff and trusting diff to create tree\n>   >   NM)  Passing tree and generating diff versus local tree for review\n> \n> Well, I guess *you* didn't read carefully.  I also spoke about the\n> value of passing around triples: ancestry, diff, and tree.  The\n> question is about linking signatures to things that humans can\n> reasonably *intend* and be reasonably held accountable for, hence one\n> of the values of signed diffs.  (I cited other practical reasons to\n> value signed diffs and use them in specific ways, too.)\n\nI know that you mentioned other things.  That doesn't invalidate that Noel\nwas talking about your starting point description of how git works and\nsuggesting that it isn't how git actually works.  The relevance of your\nother points depends upon having the base model correct.\n\nYou can argue that glass houses are inherently brittle, but why should I\ncare if mine is already made of bricks instead of glass?  If the model\nagainst which you are arguing is not the model which is used by git, then\nthe model isn't a relevant basis for claiming problems with git.\n\n\n"},{"id":"2189","messageId":"Pine.LNX.4.21.0504291706400.30848-100000@iabervon.org","threadId":"317","inReplyTo":"Pine.LNX.4.62.0504291333550.7439@qynat.qvtvafvgr.pbz","subject":"Re: git network protocol","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-04-29T21:15:01Z","receivedAt":"2005-04-29T21:15:01Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 29 Apr 2005, David Lang wrote:\n\n> would it make sense for the network git protocol to be something along the \n> lines of\n> \n> client contacts server and sends\n> the tag you want to sync with (defaults to head)\n> the local index file\n\nActually, you really want to have a bidirectional interaction, where the\nclient first fetches the info to determine where to start, and then goes\nthrough the reachable space, asking for anything it doesn't already have.\n\n(In the long run, we want to keep track of some things we already have all\nof, or know we're missing, etc., so the receiver side doesn't have to\nlook over its whole tree.)\n\ngit already includes two versions of this protocol; the first runs against\na static HTTP server, and the second uses ssh to get a socket. At some\npoint, I'm going to enable these programs to read and write\n.git/refs/?/? to figure out what they're supposed to get.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"2192","messageId":"20050429212052.GD21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0504291338540.18901@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-29T21:20:52Z","receivedAt":"2005-04-29T21:20:52Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 01:49:18PM -0700, Linus Torvalds wrote:\n> \n> \n> On Fri, 29 Apr 2005, Matt Mackall wrote:\n> > \n> > The changeset log (and everything else) has an external index.\n> \n> I don't actually know exactly how the BK changeset file works, but your \n> explanation really sounds _very_ much like it.\n\nI've never used BK, but I got the impression that it was all SCCS\nunder the covers, which means adding stuff and reconstructing random\nversions is expensive (just as it is in CVS). The split between index\nand data in Mercurial is intended to address that.\n \n> I didn't want to do anything that even smelled of BK. Of course, part of\n> my reason for that is that I didn't feel comfortable with a delta model at\n> all (I wouldn't know where to start, and I hate how they always end up\n> having different rules for \"delta\"ble and \"non-delta\"ble objects).\n\nThere aren't really any such rules here. While the index contains a\nfull DAG, the deltas are done opportunistically on a linearized\n(topologically sorted) version of it. We try to make a delta against\nthe previous tip (regardless of whether or not it's the parent), and\nif that is a win, we store it.\n\n> So it sounds like it could work fine, but it in fact sounds so much like \n> the ChangeSet file that I'd personally not have done it that way. \n\nWell I originally set out to do it differently, but I decided my\ncurrent approach was the fastest route to something that actually\nworked.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2200","messageId":"86hdhputyz.fsf@speedy.lifl.fr","threadId":"317","inReplyTo":"200504292044.NAA28429@emf.net","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Denys Duchier","fromEmail":"duchier@ps.uni-sb.de","sentAt":"2005-04-29T21:57:24Z","receivedAt":"2005-04-29T21:57:24Z","isPatch":false,"sender":{"key":"duchier@ps.uni-sb.de","avatar":null},"body":"Tom Lord <lord@emf.net> writes:\n\n>   > My example had Joe downloading a remote signed tree, reviewing the changes\n>   > locally between his own trusted tree and the remote tree, \n>\n> In the real world, that \"review\" step is the weak link.  When it goes\n> wrong, the first step is to make sure we are reviewing a tree everyone\n> involved *intended* -- and it's only with signed diffs adding up to\n> that tree that we get there.\n\nHi Tom,\n\nI hope I am not speaking out of turn or misinterpreting issues beyond my grasp,\nbut my perception of git is that when you sign a commit, you guarantee that this\nis indeed the next step in the chronology of your own branch.  It's not about\ndiffs; it's about a singular brachial chronology - of course, additional\ninformation may be recorded about topological antecedents, but that's not what\nthe signature is about.  The diff from that chronology can easily be generated\nand scrutinized by anyone, and imported or not into another branch.\n\nCheers,\n\n-- \nDr. Denys Duchier - IRI & LIFL - CNRS, Lille, France\nAIM: duchierdenys\n"},{"id":"2210","messageId":"20050429223052.GD28540@dspnet.fr.eu.org","threadId":"317","inReplyTo":"20050429201957.GJ17379@opteron.random","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Olivier Galibert","fromEmail":"galibert@pobox.com","sentAt":"2005-04-29T22:30:52Z","receivedAt":"2005-04-29T22:30:52Z","isPatch":false,"sender":{"key":"galibert@pobox.com","avatar":null},"body":"On Fri, Apr 29, 2005 at 10:19:57PM +0200, Andrea Arcangeli wrote:\n> such a system might fall apart under load, converting on the fly from\n> git to network-optimized format sound quite expensive operation, even\n> ignorign the initial decompression of the payload.\n\nNothing a little caching can't solve.  Given that git's objects are\nimmutable caching is especially easy to do, you can have the delta\nreference indexes in the filename.\n\n  OG.\n\n"},{"id":"2212","messageId":"20050429224742.GN17379@opteron.random","threadId":"317","inReplyTo":"20050429223052.GD28540@dspnet.fr.eu.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrea Arcangeli","fromEmail":"andrea@suse.de","sentAt":"2005-04-29T22:47:42Z","receivedAt":"2005-04-29T22:47:42Z","isPatch":false,"sender":{"key":"andrea@suse.de","avatar":null},"body":"On Sat, Apr 30, 2005 at 12:30:52AM +0200, Olivier Galibert wrote:\n> Nothing a little caching can't solve.  Given that git's objects are\n> immutable caching is especially easy to do, you can have the delta\n> reference indexes in the filename.\n\nRather than creating delta reference indexes we can as well use\nmercurial that uses them as primary storage.\n\ngit is the _storage_ filesystem, if we can't use it but we've to create\nanother different representation to do efficient network download, we\ncan as well the more efficient representation instead of git, that's\nwhat mercurial does AFIK.\n"},{"id":"2230","messageId":"20050430025211.GP17379@opteron.random","threadId":"317","inReplyTo":"20050429203959.GC21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrea Arcangeli","fromEmail":"andrea@suse.de","sentAt":"2005-04-30T02:52:11Z","receivedAt":"2005-04-30T02:52:11Z","isPatch":false,"sender":{"key":"andrea@suse.de","avatar":null},"body":"On Fri, Apr 29, 2005 at 01:39:59PM -0700, Matt Mackall wrote:\n> Mercurial is ammenable to rsync provided you devote a read-only\n> repository to it on the client side. In other words, you rsync from\n> kernel.org/mercurial/linus to local/linus and then you merge from\n> local/linus to your own branch. Mercurial's hashing hierarchy is\n> similar to git's (and Monotone's), so you can sign a single hash of\n> the tree as well.\n\nOk fine. It's also interesting how you already enabled partial transfers\nthrough http.\n\nPlease apply this patch so it doesn't fail on my setup ;)\n\n--- mercurial-0.4b/hg.~1~\t2005-04-29 02:52:52.000000000 +0200\n+++ mercurial-0.4b/hg\t2005-04-30 00:53:02.000000000 +0200\n@@ -1,4 +1,4 @@\n-#!/usr/bin/python\n+#!/usr/bin/env python\n #\n # mercurial - a minimal scalable distributed SCM\n # v0.4b \"oedipa maas\"\n\nOn a bit more technical side, one thing I'm wondering about is the\ncompression. If I change mercurial like this:\n\n--- revlog.py.~1~\t2005-04-29 01:33:14.000000000 +0200\n+++ revlog.py\t2005-04-30 03:54:12.000000000 +0200\n@@ -11,9 +11,11 @@\n import zlib, struct, mdiff, sha, binascii, os, tempfile\n \n def compress(text):\n+    return text\n     return zlib.compress(text)\n \n def decompress(bin):\n+    return text\n     return zlib.decompress(bin)\n \n def hash(text):\n\n\nthe .hg directory sizes changes from 167M to 302M _BUT_ the _compressed_\nsize of the .hg directory (i.e. like in a full network transfer with\nrsync -z or a tar.gz backup) changes from 55M to 38M:\n\nandrea@opteron:~/devel/kernel> du -sm hg-orig hg-aa hg-orig.tar.bz2 hg-aa.tar.bz2 \n167     hg-orig\n302     hg-aa\n55      hg-orig.tar.bz2\n38      hg-aa.tar.bz2\n^^^^^^^^^^^^^^^^^^^^^ 38M backup and network transfer is what I want\n\nSo I don't really see an huge benefit in compression, other than to\nslowdown the checkins measurably [i.e. what Linus doesn't want] (the\ntime of compression is a lot higher than the time of python runtime during\ncheckin, so it's hard to believe your 100% boost with psyco in the hg file,\nsometime psyco doesn't make any difference infact, I'd rather prefer people to\nwork on the real thing of generating native bytecode at compile time, rather\nthan at runtime, like some haskell compiler can do).\n\nmercurial is already good at decreasing the entropy by using an efficient\nstorage format, it doesn't need to cheat by putting compression on each blob\nthat can only leads to bad ratios when doing backups and while transferring\nmore than one blob through the network.\n\nSo I suggest to try disabling compression optionally, perhaps it'll be even\nfaster than git in the initial checkin that way! No need of compressing or\ndecompressing anything with mercurial (unlike with git that would explode\nwithout control w/o compression).\n\nMy time measurements follows:\n\nw/o compression:\n\n\t9.52user 73.11system 1:30.49elapsed 91%CPU (0avgtext+0avgdata 0maxresident)k\n\t^\n\t0inputs+0outputs (0major+80109minor)pagefaults 0swaps\n\nw/ compression (i.e. official package):\n\n\t26.78user 75.90system 1:44.87elapsed 97%CPU (0avgtext+0avgdata 0maxresident)k\n\t^^\n\t0inputs+0outputs (0major+484522minor)pagefaults 0swaps\n\nThe user time is by far the most important reliable number here, 17 seconds of\ndifference wasted in compression time. The 1:30 time w/o cache didn't fit\ncompletely in cache, but still it was faster (only 14 sec faster instead of 17\nsec faster due some minor I/O that happened due the larger pagecache size that\nwas recycled a bit).\n\nWithout compression the time is 90% system time, very little time is spent in\nuserland, and of course I made sure very little time is spent in I/O.  vmstat\nlooks like this:\n\n 1  0 107008  61320  31964 546076    0    0     0  7860 1095  1147  5 46 41  9\n 1  0 107008  59336  31972 547564    0    0     0     0 1074  1093  6 46 48  0\n 1  0 107008  57544  31988 549044    0    0     0     0 1095  1239  5 45 50  0\n 1  0 107008  55304  32000 550936    0    0     0     0 1064  1041  5 46 50  0\n 1  0 107008  53384  32012 552488    0    0     0     0 1080  1081  5 45 50  0\n 1  0 107008  51592  32032 553964    0    0     0  8260 1087  1018  3 48 40 10\n 1  0 107008  49144  32040 555180    0    0     0     0 1099  1090  5 48 47  0\n 1  0 107008  47432  32048 556532    0    0     0     0 1086  1014  4 46 50  0\n 1  0 107008  45632  32060 557676    0    0     0     0 1102  1073  4 47 49  0\n 2  0 107008  44032  32068 558892    0    0     0     0 1088  1044  4 46 49  0\n 1  0 107008  42864  32116 560204    0    0     8  6672 1136  1265  5 47 41  7\n 1  0 107008  41008  32124 561420    0    0     0   484 1182  1078  5 43 49  3\n\n(5 = user, 43 system, 49 idle is the second cpu doing nothing, 3 is io-wait\ntime)\n\nWhile with compression (default) the user time goes up, a lot higher than\nwhat python wasted in the above trace:\n\n 1  0 107008 282688  26756 346396    0    0     0 12064 1122   997 20 32 38 12\n 1  0 107008 279936  26776 348552    0    0     0     0 1074   938 15 36 50  0\n 1  0 107008 277296  26780 350656    0    0     0     0 1087  1070 15 36 50  0\n 1  0 107008 274672  26796 352816    0    0     4    28 1060  1021 15 36 49  1\n 1  0 107008 272176  26824 354828    0    0     0    52 1092  1082 19 32 50  0\n 1  1 107008 269616  26844 356780    0    0     0 10856 1106  1019 16 36 36 13\n 1  0 107008 267312  26864 358936    0    0     0     4 1081  1068 27 24 49  0\n 1  0 107008 265072  26876 360760    0    0     0     0 1068  1073 23 27 50  0\n 1  0 107008 263024  26888 362516    0    0     0  2764 1224  1457 18 29 49  5\n 1  0 107008 260928  26900 364408    0    0     0    20 1060   969 21 29 50  0\n 1  0 107008 258752  26928 366216    0    0     0  9768 1098   916 18 32 37 12\n 1  0 107008 256640  26940 367972    0    0     0     0 1058  1093 13 37 50  0\n 1  0 107008 254384  26952 369796    0    0     0     0 1089  1259 18 33 50  0\n 1  0 107008 252016  26972 371680    0    0     0     0 1063  1040 29 22 50  0\n\nThe difference in time may be even higher than the above since the pass w/o\ncompression may have done some read I/O too sometime,  since I've 1G and\nit didn't fit completely in cache (I made sure the second pass was completely\nfrom cache instead, by running twice and verifying the free memory never went\nlower than 100M).\n\nHttp is not intended for maximal efficiency, it's there just to make\nlife easy. special protocol with zlib is required for maximum\nefficiency.\n\nMy suggestion is to leave compression optional by storing a bit in the .hg\ndirectory (a global bit) to know if the blob is compressed or not... so if\npeople have problems and runs out of pagecache they can use it, but Linus\ndefinitely must keep compression off to run the checkin at max speed (i.e.\npotentially much faster than git, especially on big files). The cost of one\nbranch in python to know if compression is enabled or not shouldn't be an issue.\nCompression should be a parameter to \"hg init\", perhaps hg -z, so in the future\nwe can add -j to add bzip2 too.\n\nYou also should move the .py into a hg directory, so that they won't\npollute the site-packages.\n\nMatt, very great work with mercurial. The small mercurial size is striking, 1037\nlines total excluding the classes you rightfully shared from other python\nprojects. That's less than 1/7 of the size of cogito (ok perhaps it's not as\nmature as cogito but anyway). Great choice for the language too (but\nperhaps I'm biased ;).\n"},{"id":"2258","messageId":"20050430152014.GI21897@waste.org","threadId":"317","inReplyTo":"20050430025211.GP17379@opteron.random","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-04-30T15:20:15Z","receivedAt":"2005-04-30T15:20:15Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Sat, Apr 30, 2005 at 04:52:11AM +0200, Andrea Arcangeli wrote:\n> On Fri, Apr 29, 2005 at 01:39:59PM -0700, Matt Mackall wrote:\n> > Mercurial is ammenable to rsync provided you devote a read-only\n> > repository to it on the client side. In other words, you rsync from\n> > kernel.org/mercurial/linus to local/linus and then you merge from\n> > local/linus to your own branch. Mercurial's hashing hierarchy is\n> > similar to git's (and Monotone's), so you can sign a single hash of\n> > the tree as well.\n> \n> Ok fine. It's also interesting how you already enabled partial transfers\n> through http.\n> \n> Please apply this patch so it doesn't fail on my setup ;)\n> \n> --- mercurial-0.4b/hg.~1~\t2005-04-29 02:52:52.000000000 +0200\n> +++ mercurial-0.4b/hg\t2005-04-30 00:53:02.000000000 +0200\n> @@ -1,4 +1,4 @@\n> -#!/usr/bin/python\n> +#!/usr/bin/env python\n\nDone.\n\n> On a bit more technical side, one thing I'm wondering about is the\n> compression. If I change mercurial like this:\n> \n> --- revlog.py.~1~\t2005-04-29 01:33:14.000000000 +0200\n> +++ revlog.py\t2005-04-30 03:54:12.000000000 +0200\n> @@ -11,9 +11,11 @@\n>  import zlib, struct, mdiff, sha, binascii, os, tempfile\n>  \n>  def compress(text):\n> +    return text\n>      return zlib.compress(text)\n>  \n>  def decompress(bin):\n> +    return text\n>      return zlib.decompress(bin)\n>  \n>  def hash(text):\n> \n> \n> the .hg directory sizes changes from 167M to 302M _BUT_ the _compressed_\n> size of the .hg directory (i.e. like in a full network transfer with\n> rsync -z or a tar.gz backup) changes from 55M to 38M:\n> \n> andrea@opteron:~/devel/kernel> du -sm hg-orig hg-aa hg-orig.tar.bz2 hg-aa.tar.bz2 \n> 167     hg-orig\n> 302     hg-aa\n> 55      hg-orig.tar.bz2\n> 38      hg-aa.tar.bz2\n> ^^^^^^^^^^^^^^^^^^^^^ 38M backup and network transfer is what I want\n> \n> So I don't really see an huge benefit in compression, other than to\n> slowdown the checkins measurably [i.e. what Linus doesn't want] (the\n> time of compression is a lot higher than the time of python runtime during\n> checkin, so it's hard to believe your 100% boost with psyco in the hg file,\n> sometime psyco doesn't make any difference infact, I'd rather prefer people to\n> work on the real thing of generating native bytecode at compile time, rather\n> than at runtime, like some haskell compiler can do).\n\nMost of that psyco speed up is accelerating subsequent diffs in\ndifflib, which you probably didn't hit yet.\n\n> mercurial is already good at decreasing the entropy by using an efficient\n> storage format, it doesn't need to cheat by putting compression on each blob\n> that can only leads to bad ratios when doing backups and while transferring\n> more than one blob through the network.\n> \n> So I suggest to try disabling compression optionally, perhaps it'll be even\n> faster than git in the initial checkin that way! No need of compressing or\n> decompressing anything with mercurial (unlike with git that would explode\n> without control w/o compression).\n\nI can make it some sort of environment variable, sure. I think the\nspeed is already in a domain where it's not a big deal though. There\nare other things to do first, like unifying the merge/commit/update\ncode.\n\n> Http is not intended for maximal efficiency, it's there just to make\n> life easy. special protocol with zlib is required for maximum\n> efficiency.\n\nYeah, I've got a plan here.\n\n> You also should move the .py into a hg directory, so that they won't\n> pollute the site-packages.\n\nYep, I'm rather new to actually packaging my Python hacks.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2262","messageId":"20050430163724.GE20146@opteron.random","threadId":"317","inReplyTo":"20050430152014.GI21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrea Arcangeli","fromEmail":"andrea@suse.de","sentAt":"2005-04-30T16:37:24Z","receivedAt":"2005-04-30T16:37:24Z","isPatch":false,"sender":{"key":"andrea@suse.de","avatar":null},"body":"On Sat, Apr 30, 2005 at 08:20:15AM -0700, Matt Mackall wrote:\n> Most of that psyco speed up is accelerating subsequent diffs in\n> difflib, which you probably didn't hit yet.\n\nCorrect. Plus I've a 64bit python so I can't use psyco anyway.\n\n> I can make it some sort of environment variable, sure. I think the\n> speed is already in a domain where it's not a big deal though. There\n\nNo big deal of course, I mentioned it just because it was by far the\nmost CPU userland intensive operation during checkin. Perhaps doing less\nvfs syscalls would improve checkin time too, but I'm unsure if that's\neasily feasible (while disabling compression was certainly easy ;)\n\n> Yep, I'm rather new to actually packaging my Python hacks.\n\nI sent you by private email a modified package that gets that right.\n\nThanks!\n"},{"id":"2366","messageId":"42764C0C.8030604@tmr.com","threadId":"317","inReplyTo":"20050430025211.GP17379@opteron.random","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-05-02T15:49:32Z","receivedAt":"2005-05-02T15:49:32Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Andrea Arcangeli wrote:\n> On Fri, Apr 29, 2005 at 01:39:59PM -0700, Matt Mackall wrote:\n> \n>>Mercurial is ammenable to rsync provided you devote a read-only\n>>repository to it on the client side. In other words, you rsync from\n>>kernel.org/mercurial/linus to local/linus and then you merge from\n>>local/linus to your own branch. Mercurial's hashing hierarchy is\n>>similar to git's (and Monotone's), so you can sign a single hash of\n>>the tree as well.\n> \n> \n> Ok fine. It's also interesting how you already enabled partial transfers\n> through http.\n> \n> Please apply this patch so it doesn't fail on my setup ;)\n> \n> --- mercurial-0.4b/hg.~1~\t2005-04-29 02:52:52.000000000 +0200\n> +++ mercurial-0.4b/hg\t2005-04-30 00:53:02.000000000 +0200\n> @@ -1,4 +1,4 @@\n> -#!/usr/bin/python\n> +#!/usr/bin/env python\n>  #\n>  # mercurial - a minimal scalable distributed SCM\n>  # v0.4b \"oedipa maas\"\n\nCould you explain why this is necessary or desirable? I looked at what \nenv does, and I am missing the point of duplicating bash normal \nbehaviour regarding definition of per-process environment entries.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n\n"},{"id":"2386","messageId":"427650E7.2000802@tmr.com","threadId":"317","inReplyTo":"20050429165232.GV21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-05-02T16:10:15Z","receivedAt":"2005-05-02T16:10:15Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Matt Mackall wrote:\n> On Fri, Apr 29, 2005 at 11:18:20AM -0400, Morten Welinder wrote:\n> \n>>>I had three design goals. \"disk space\" wasn't one of them\n>>\n>>And, if at some point it should become an issue, it's fixable. Since\n>>access to objects is fairly centralized and since they are\n>>immutable, it would be quite simple to move an arbitrary selection\n>>of the objects into some other storage form which could take\n>>similarities between objects into account.\n> \n> \n> This is not a fix, this is a band-aid. A fix is fitting all the data\n> in 10 times less space without sacrificing too much performance.\n> \n> \n>>So disk space and its cousin number-of-files are both when-and-if\n>>problems. And not scary ones at that.\n> \n> \n> But its sibling bandwidth _is_ a problem. The delta between 2.6.10 and\n> 2.6.11 in git terms will be much larger than a _full kernel tarball_.\n> Simply checking in patch-2.6.11 on top of 2.6.10 as a single changeset\n> takes 41M. Break that into a thousand overlapping deltas (ie the way\n> it is actually done) and it will be much larger.\n> \nAt this level of performance I would say it doesn't matter. If a full \ncheckin take two minutes or three minutes doesn't concern me, because \nI'm not going to sit and watch it, I'm going to read LKML or write my \nbeer blog in another window. I would care about two vs. three hours, but \nminutes are too long to wait and too short to care.\n\nNow look at pulling 41MB over a T1 link. All of a sudden I care bigtime! \nI want very much to use my bandwidth for other things, I don't want 41MB \nadded to my backup, etc. Disk space is cheap, but unless you ignore \nbackups and have an OC3 or so, these numbers are large enough to be \nirritating. Not a huge issue, just one of those \"piss me off every time \nI do it\" things.\n\nIf there is a functional reason to use git, something Mercurial doesn't \ndo, then developers will and should use git. But the associated hassles \nwith large change size, rather than the absolute size, are worth \nconsidering.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n\n"},{"id":"2369","messageId":"200505021614.j42GEufG008441@turing-police.cc.vt.edu","threadId":"317","inReplyTo":"42764C0C.8030604@tmr.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"","fromEmail":"valdis.kletnieks@vt.edu","sentAt":"2005-05-02T16:14:56Z","receivedAt":"2005-05-02T16:14:56Z","isPatch":false,"sender":{"key":"valdis.kletnieks@vt.edu","avatar":null},"body":"On Mon, 02 May 2005 11:49:32 EDT, Bill Davidsen said:\n> Andrea Arcangeli wrote:\n> > On Fri, Apr 29, 2005 at 01:39:59PM -0700, Matt Mackall wrote:\n\n> > -#!/usr/bin/python\n> > +#!/usr/bin/env python\n> >  #\n> >  # mercurial - a minimal scalable distributed SCM\n> >  # v0.4b \"oedipa maas\"\n> \n> Could you explain why this is necessary or desirable? I looked at what \n> env does, and I am missing the point of duplicating bash normal \n> behaviour regarding definition of per-process environment entries.\n\nMost likely, his python lives elsewhere than /usr/bin, and the 'env' call\nresults in causing a walk across $PATH to find it....\n"},{"id":"2370","messageId":"42765212.8030605@tmr.com","threadId":"317","inReplyTo":"2944.10.10.10.24.1114802002.squirrel@linux1","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-05-02T16:15:14Z","receivedAt":"2005-05-02T16:15:14Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Sean wrote:\n> On Fri, April 29, 2005 2:54 pm, Tom Lord said:\n> \n> \n>>The process should not rely on the security of every developer's\n>>machine.  The process should not rely on simply trusting quality\n>>contributors by reputation (e.g., most cons begin by establishing\n>>trust and continue by relying inappropriately on\n>>trust-without-verification).  This relates to why Linus'\n>>self-advertised process should be raising yellow and red cards all\n>>over the place: either he is wasting a huge amount of his own time and\n>>should be largely replaced by an automated patch queue manager, or he\n>>is being trusted to do more than is humanly possible.\n>>\n> \n> \n> Ahh, you don't believe in the development model that has produced Linux! \n> Personally I do believe in it, so much so that I question the value of\n> signatures at the changeset level.  To me it doesn't matter where the code\n> came from just so long as it works.\n\nLawyers must love you... That approach doesn't work in court.\n\nRelated: look at the new software patent law, it ignores the existing \nlaw, judge and jury, and lets MS avoid paying the judgement for a suit \nit already lost.\n\nSee Computerworld etc for details.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n"},{"id":"2368","messageId":"20050502161722.GN20146@opteron.random","threadId":"317","inReplyTo":"42764C0C.8030604@tmr.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Andrea Arcangeli","fromEmail":"andrea@suse.de","sentAt":"2005-05-02T16:17:22Z","receivedAt":"2005-05-02T16:17:22Z","isPatch":false,"sender":{"key":"andrea@suse.de","avatar":null},"body":"On Mon, May 02, 2005 at 11:49:32AM -0400, Bill Davidsen wrote:\n> Could you explain why this is necessary or desirable? I looked at what \n\nThis is necessary here because of this:\n\nandrea@opteron:~> which python\n/home/andrea/bin/x86_64/python/bin/python\n\nOf course I've /home/andrea/bin/x86_64/python/bin in the path before\n/usr/bin.\n\nThe generally accepted way to start it is through env, other scripts in\nmercurial were already getting that right too so it was probably not\nintentional to hardcode it in the hg binary.\n"},{"id":"2371","messageId":"Pine.LNX.4.58.0505020921080.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"42764C0C.8030604@tmr.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-02T16:31:06Z","receivedAt":"2005-05-02T16:31:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Bill Davidsen wrote:\n> > -#!/usr/bin/python\n> > +#!/usr/bin/env python\n\n> Could you explain why this is necessary or desirable? I looked at what \n> env does, and I am missing the point of duplicating bash normal \n> behaviour regarding definition of per-process environment entries.\n\nIt's not about environment.\n\nIt's about the fact that many people have things like python in\n/usr/local/bin/python, because they compiled it themselves or similar.\n\nPretty much the only path you can _really_ depend on for #! stuff is \n/bin/sh.\n\nAny system that doesn't have /bin/sh is so fucked up that it's not worth\nworrying about. Anything else can be in /bin, /usr/bin or /usr/local/bin\n(and sometimes other strange places).\n\nThat said, I think the /usr/bin/env trick is stupid too. It may be more \nportable for various Linux distributions, but if you want _true_ \nportability, you use /bin/sh, and you do something like\n\n\t#!/bin/sh\n\texec perl perlscript.pl \"$@\"\n\ninstead.\n\t\tLinus\n"},{"id":"2375","messageId":"20050502171802.GA28045@nevyn.them.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505020921080.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Daniel Jacobowitz","fromEmail":"dan@debian.org","sentAt":"2005-05-02T17:18:02Z","receivedAt":"2005-05-02T17:18:02Z","isPatch":false,"sender":{"key":"dan@debian.org","avatar":null},"body":"On Mon, May 02, 2005 at 09:31:06AM -0700, Linus Torvalds wrote:\n> It's not about environment.\n> \n> It's about the fact that many people have things like python in\n> /usr/local/bin/python, because they compiled it themselves or similar.\n> \n> Pretty much the only path you can _really_ depend on for #! stuff is \n> /bin/sh.\n> \n> Any system that doesn't have /bin/sh is so fucked up that it's not worth\n> worrying about. Anything else can be in /bin, /usr/bin or /usr/local/bin\n> (and sometimes other strange places).\n> \n> That said, I think the /usr/bin/env trick is stupid too. It may be more \n> portable for various Linux distributions, but if you want _true_ \n> portability, you use /bin/sh, and you do something like\n> \n> \t#!/bin/sh\n> \texec perl perlscript.pl \"$@\"\n> \n> instead.\n\nDo you know any vaguely Unix-like system where #!/usr/bin/env does not\nwork?  I don't; I've used it on Solaris, HP-UX, OSF/1...\n\n-- \nDaniel Jacobowitz\nCodeSourcery, LLC\n"},{"id":"2376","messageId":"20050502172012.GD11726@mythryan2.michonline.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505020921080.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2005-05-02T17:20:12Z","receivedAt":"2005-05-02T17:20:12Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"On Mon, May 02, 2005 at 09:31:06AM -0700, Linus Torvalds wrote:\n> That said, I think the /usr/bin/env trick is stupid too. It may be more \n> portable for various Linux distributions, but if you want _true_ \n> portability, you use /bin/sh, and you do something like\n> \n> \t#!/bin/sh\n> \texec perl perlscript.pl \"$@\"\n\t\tif 0;\n\nYou don't really want Perl to get itself into an exec loop.\n\n-- \n\nRyan Anderson\n  sometimes Pug Majere\n"},{"id":"2378","messageId":"Pine.LNX.4.58.0505021028540.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"20050502172012.GD11726@mythryan2.michonline.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-02T17:31:04Z","receivedAt":"2005-05-02T17:31:04Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Ryan Anderson wrote:\n>\n> On Mon, May 02, 2005 at 09:31:06AM -0700, Linus Torvalds wrote:\n> > That said, I think the /usr/bin/env trick is stupid too. It may be more \n> > portable for various Linux distributions, but if you want _true_ \n> > portability, you use /bin/sh, and you do something like\n> > \n> > \t#!/bin/sh\n> > \texec perl perlscript.pl \"$@\"\n> \t\tif 0;\n> \n> You don't really want Perl to get itself into an exec loop.\n\nThis would _not_ be \"perlscript.pl\" itself. This is the shell-script, and \nit's not called \".pl\".\n\nIn other words, you'd put this as ~/bin/cg-xxxx, and then \"perlscript.pl\" \nwouldn't be in the path at all, it would be in some separate install \ndirectory.\n\nBut hey, if people want to be safe for bad installations, add the extra \nline. Shell won't care ;)\n\n\t\tLinus\n"},{"id":"2379","messageId":"Pine.LNX.4.58.0505021031070.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"20050502171802.GA28045@nevyn.them.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-02T17:32:38Z","receivedAt":"2005-05-02T17:32:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Daniel Jacobowitz wrote:\n> \n> Do you know any vaguely Unix-like system where #!/usr/bin/env does not\n> work?  I don't; I've used it on Solaris, HP-UX, OSF/1...\n\nI've used unixes where \"#!\" didn't work.\n\nThings like bash still have support for such unixes, I think - you can\ntell them to parse the #! line themselves, to make it appear to do the \nright thing.\n\nAre these common? Hell no. But they definitely existed.\n\n\t\tLinus\n"},{"id":"2385","messageId":"20050502201748.0557cc0e.froese@gmx.de","threadId":"317","inReplyTo":"20050502171802.GA28045@nevyn.them.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Edgar Toernig","fromEmail":"froese@gmx.de","sentAt":"2005-05-02T18:17:48Z","receivedAt":"2005-05-02T18:17:48Z","isPatch":false,"sender":{"key":"froese@gmx.de","avatar":null},"body":"Daniel Jacobowitz wrote:\n>\n> Do you know any vaguely Unix-like system where #!/usr/bin/env does not\n> work?  I don't; I've used it on Solaris, HP-UX, OSF/1...\n\nOn old System-Vs (that includes *caugh* SCO-Unix) env was in /bin not\n/usr/bin.\n\nCiao, ET.\n"},{"id":"2389","messageId":"1785.10.10.10.24.1115060548.squirrel@linux1","threadId":"317","inReplyTo":"427650E7.2000802@tmr.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2005-05-02T19:02:28Z","receivedAt":"2005-05-02T19:02:28Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Mon, May 2, 2005 12:10 pm, Bill Davidsen said:\n\n> Now look at pulling 41MB over a T1 link. All of a sudden I care bigtime!\n> I want very much to use my bandwidth for other things, I don't want 41MB\n> added to my backup, etc. Disk space is cheap, but unless you ignore\n> backups and have an OC3 or so, these numbers are large enough to be\n> irritating. Not a huge issue, just one of those \"piss me off every time\n> I do it\" things.\n\nThat 41MB or lets say 200MB is spread over several months between\nreleases.   Pulling once a day from the git public repository, makes this\nbarely noticeable.  In the future there may be optimized protocols to\nhandle this more efficiently.\n\nYou bring up a good point about backups though.  Eventually it might be\nnice to have a utility that exports/imports a git repository in a flat\nfile using deltas rather than snapshots.   Such an export format would\nmake backups and tarballs cheaper.\n\nSean\n\n\n"},{"id":"2399","messageId":"20050502205418.GA12409@mars.ravnborg.org","threadId":"317","inReplyTo":"20050502171802.GA28045@nevyn.them.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Sam Ravnborg","fromEmail":"sam@ravnborg.org","sentAt":"2005-05-02T20:54:18Z","receivedAt":"2005-05-02T20:54:18Z","isPatch":false,"sender":{"key":"sam@ravnborg.org","avatar":"https://gravatar.com/avatar/168a912606ed0742d840bb365e3cc21db390c36531a58341dc7a069cc1f15f62?d=mp&s=160"},"body":"On Mon, May 02, 2005 at 01:18:02PM -0400, Daniel Jacobowitz wrote:\n> > \t#!/bin/sh\n> > \texec perl perlscript.pl \"$@\"\n> > \n> > instead.\n> \n> Do you know any vaguely Unix-like system where #!/usr/bin/env does not\n> work?  I don't; I've used it on Solaris, HP-UX, OSF/1...\n\nI had to pull out a call to env from kbuild due to strange errors in\nsome mandrake? based system.\nI never tracked it down fully at that time, I just realised that two\ndifferent programs named env was present, and the less common one made\nthe linux kernel build fail. env was not called with any path in that\nexample so that may have cured it.\n\n\tSam\n"},{"id":"2402","messageId":"F9443EC8-A8F7-47D7-AC64-AEC476E7223F@mac.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505020921080.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Kyle Moffett","fromEmail":"mrmacman_g4@mac.com","sentAt":"2005-05-02T21:17:32Z","receivedAt":"2005-05-02T21:17:32Z","isPatch":false,"sender":{"key":"mrmacman_g4@mac.com","avatar":null},"body":"On May 2, 2005, at 12:31:06, Linus Torvalds wrote:\n> That said, I think the /usr/bin/env trick is stupid too. It may be  \n> more\n> portable for various Linux distributions, but if you want _true_\n> portability, you use /bin/sh, and you do something like\n>\n>     #!/bin/sh\n>     exec perl perlscript.pl \"$@\"\n\nOooh, I can one-up that hack with this evil from perlrun(1):\n\n#!/bin/sh -- # -*- perl -*- -W -T\neval 'exec perl -wS $0 ${1+\"$@\"}'\n     if 0;\n# PERL SCRIPT HERE\n\nDescription:\nPerl ignores the eval($string) because of the \"if 0\" in the  \nstatement. The\nshell sees the statement end at the newline, and executes it faithfully.\nThe end result is that the preferred Perl gets the script.  I don't know\nPython, so I don't know if such a trick exists there.\n\nCheers,\nKyle Moffett\n\n-----BEGIN GEEK CODE BLOCK-----\nVersion: 3.12\nGCM/CS/IT/U d- s++: a18 C++++>$ UB/L/X/*++++(+)>$ P+++(++++)>$\nL++++(+++) E W++(+) N+++(++) o? K? w--- O? M++ V? PS+() PE+(-) Y+\nPGP+++ t+(+++) 5 X R? tv-(--) b++++(++) DI+ D+ G e->++++$ h!*()>++$  \nr  !y?(-)\n------END GEEK CODE BLOCK------\n\n\n\n"},{"id":"2403","messageId":"Pine.LNX.4.58.0505021457060.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"427650E7.2000802@tmr.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-02T22:02:16Z","receivedAt":"2005-05-02T22:02:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Bill Davidsen wrote:\n> \n> If there is a functional reason to use git, something Mercurial doesn't \n> do, then developers will and should use git. But the associated hassles \n> with large change size, rather than the absolute size, are worth \n> considering.\n\nNote that we discussed this early on, and the issues with full-file \nhandling haven't changed. It does actually have real functional \nadvantages:\n\n - you can share the objects freely between different trees, never \n   worrying about one tree corrupting another trees object by mistake.\n - you can drop old objects.\n\ndelta models very fundamentally don't support this. \n\nFor example, a simple tree re-linker will work on any mirror site, and\nwork reliably, even if I end up uploading new objects with some tool that\ndoesn't know to break hardlinks etc. That can easily be much more than a\n10x win for a git repository site (imagine something like bkbits.net, but\ngot git).\n\nWhether it is a huge deal or not, I don't know. I do know that the big \ndeal to me is just the simplicity of the git object models. It makes me \ntrust it, even in the presense of inevitable bugs. It's a very safe model, \nand right now safe is good.\n\n\t\tLinus\n"},{"id":"2408","messageId":"20050502223002.GP21897@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505021457060.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-02T22:30:02Z","receivedAt":"2005-05-02T22:30:02Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Mon, May 02, 2005 at 03:02:16PM -0700, Linus Torvalds wrote:\n> \n> \n> On Mon, 2 May 2005, Bill Davidsen wrote:\n> > \n> > If there is a functional reason to use git, something Mercurial doesn't \n> > do, then developers will and should use git. But the associated hassles \n> > with large change size, rather than the absolute size, are worth \n> > considering.\n> \n> Note that we discussed this early on, and the issues with full-file \n> handling haven't changed. It does actually have real functional \n> advantages:\n> \n>  - you can share the objects freely between different trees, never \n>    worrying about one tree corrupting another trees object by mistake.\n\nNot sure if this is terribly useful. It just makes it harder to pull\nthe subset you're interested in.\n\n>  - you can drop old objects.\n\nYou can't drop old objects without dropping all the changesets that\nrefer to them or otherwise being prepared to deal with the broken\nlinks.\n\n> delta models very fundamentally don't support this. \n\nThe latter can be done in a pretty straightforward manner in mercurial\nwith one pass over the data. But I have a goal to make keeping the\nwhole history cheap enough that no one balks at it.\n\n> For example, a simple tree re-linker will work on any mirror site, and\n> work reliably, even if I end up uploading new objects with some tool that\n> doesn't know to break hardlinks etc. That can easily be much more than a\n> 10x win for a git repository site (imagine something like bkbits.net, but\n> got git).\n\nWhat is a tree re-linker? Finds duplicate files and hard-links them?\nOk, that makes some sense. But it's a win on one machine and a lose\neverywhere else.\n\n> Whether it is a huge deal or not, I don't know. I do know that the big \n> deal to me is just the simplicity of the git object models. It makes me \n> trust it, even in the presense of inevitable bugs. It's a very safe model, \n> and right now safe is good.\n\nI've added an \"hg verify\" command to Mercurial. It doesn't attempt to\nfix anything up yet, but it can catch a couple things that git\nprobably can't (like file revisions that aren't owned by any\nchangeset), namely because there's more metadata around to look at.\n\nI'll probably post an updated version tomorrow, I'm beginning to work\non a git2hg script.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2411","messageId":"Pine.LNX.4.58.0505021540070.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"20050502223002.GP21897@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-02T22:49:49Z","receivedAt":"2005-05-02T22:49:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Matt Mackall wrote:\n> > \n> >  - you can share the objects freely between different trees, never \n> >    worrying about one tree corrupting another trees object by mistake.\n> \n> Not sure if this is terribly useful. It just makes it harder to pull\n> the subset you're interested in.\n\nYou don't have to share things in a single subdirectory. Symlinks and \nhardlinks work fine, as do actual filesystem tricks ;)\n\n> >  - you can drop old objects.\n> \n> You can't drop old objects without dropping all the changesets that\n> refer to them or otherwise being prepared to deal with the broken\n> links.\n\nAbsolutely. This needs support from fsck to allow us to say \"commit xxxx \nis no longer in the tree, because we pruned it\".\n\nAlternatively (and that's the much less intrusive one), you keep all the\ncommit objects, but drop the tree and blob objects. Again, all you need \nfor this to work is just feed a list of commits to fsck, and tell it \n\"we've pruned those from the tree\", which tells fsck not to start looking \nfor the contents of those commits.\n\nSo for example, you can trivially have something that automates this: take \neach commit that is older than <x> days, add it to the \"prune list\", and \nrun fsck, and delete all objects that now show up as being unreachable \n(since fsck won't be looking at what those commits reference).\n\nI could write this up in ten minutes. It's really simple.\n\nAnd it's simple _exactly_ because we don't do deltas.\n\n> > delta models very fundamentally don't support this. \n> \n> The latter can be done in a pretty straightforward manner in mercurial\n> with one pass over the data. But I have a goal to make keeping the\n> whole history cheap enough that no one balks at it.\n\nWith delta's, you have two choices:\n\n - change all the sha1 names (ie a pruned tree would no longer be \n   compatible with a non-pruned one)\n - make the delta part not show up as part of the sha1 name (which means \n   that it's unprotected).\n\nwhich one would you have?\n\n> What is a tree re-linker? Finds duplicate files and hard-links them?\n> Ok, that makes some sense. But it's a win on one machine and a lose\n> everywhere else.\n\nWhere would it be a loss? Esepcially since with git, it's cheap (you don't \nneed to compare content to find objects to link - you can just compare \nfilename listings).\n\n> I've added an \"hg verify\" command to Mercurial. It doesn't attempt to\n> fix anything up yet, but it can catch a couple things that git\n> probably can't (like file revisions that aren't owned by any\n> changeset), namely because there's more metadata around to look at.\n\ngit-fsck-cache catches exactly those kinds of things. And since it checks\npretty much every _single_ assumption in git (which is not a lot, since\ngit doesn't have a lot of assumptions), I guarantee you that you can't\nfind any more than it does (the filename ordering is the big missing\npiece: I _still_ don't verify that trees are ordered. I've been mentioning\nit since the beginning, but I'm lazy).\n\nIn other words, your verifier can't verify anything more. It's entirely \npossible that more things can go _wrong_, since you have more indexes, so \nyour verifier will have more to check, but that's not an advantage, that's \na downside.\n\n\t\tLinus\n"},{"id":"2417","messageId":"20050503000011.GA22038@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505021540070.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-03T00:00:12Z","receivedAt":"2005-05-03T00:00:12Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Mon, May 02, 2005 at 03:49:49PM -0700, Linus Torvalds wrote:\n> > >  - you can drop old objects.\n> > \n> > You can't drop old objects without dropping all the changesets that\n> > refer to them or otherwise being prepared to deal with the broken\n> > links.\n\n[...]\n \n> I could write this up in ten minutes. It's really simple.\n\nIt's still simple in Mercurial, but more importantly Mercurial _won't\nneed it_. Dropping history is a work-around, not a feature.\n \n> > > delta models very fundamentally don't support this. \n> > \n> > The latter can be done in a pretty straightforward manner in mercurial\n> > with one pass over the data. But I have a goal to make keeping the\n> > whole history cheap enough that no one balks at it.\n> \n> With delta's, you have two choices:\n> \n>  - change all the sha1 names (ie a pruned tree would no longer be \n>    compatible with a non-pruned one)\n>  - make the delta part not show up as part of the sha1 name (which means \n>    that it's unprotected).\n> \n> which one would you have?\n\nUmm.. I am _not_ calculating the SHA of the delta itself. That'd be\nsilly.\n\nThere are an arbitrary number of ways to calculate a delta between two\nfiles. Similarly, there are an arbitrary number of ways to compress a\nfile (gzip has at least 9, not counting all the permutations of\nflush). The only sensible thing to do is store a hash of the raw text\nand check it against the fully restored text, because that's what you\ncare about being correct.\n\nIn Mercurial, deltas are just a storage detail and are effectively\ncompletely hidden from everything except the innermost part of the\nback-end. What's important is that Mercurial knows that A is a\nrevision of B in the backend and thus has enough information to\nopportunistically attempt to calculate a delta.\n\nSo if the day ever comes when I want to prune the head of a log, I\nsimply reconstruct the first version to keep, store it in a new file,\nthen append all the deltas, unmodified. And fix up the offsets in the\nindices. None of the hashes change.\n\n> > What is a tree re-linker? Finds duplicate files and hard-links them?\n> > Ok, that makes some sense. But it's a win on one machine and a lose\n> > everywhere else.\n> \n> Where would it be a loss? Esepcially since with git, it's cheap (you don't \n> need to compare content to find objects to link - you can just compare \n> filename listings).\n\nGit repositories will be 10x larger than Mercurial everywhere that\ndoesn't benefit from this linking of unrelated trees. That is, folks\nwho aren't running gitbits.net.\n\n> > I've added an \"hg verify\" command to Mercurial. It doesn't attempt to\n> > fix anything up yet, but it can catch a couple things that git\n> > probably can't (like file revisions that aren't owned by any\n> > changeset), namely because there's more metadata around to look at.\n> \n> git-fsck-cache catches exactly those kinds of things. And since it checks\n> pretty much every _single_ assumption in git (which is not a lot, since\n> git doesn't have a lot of assumptions), I guarantee you that you can't\n> find any more than it does (the filename ordering is the big missing\n> piece: I _still_ don't verify that trees are ordered. I've been mentioning\n> it since the beginning, but I'm lazy).\n> \n> In other words, your verifier can't verify anything more. It's entirely \n> possible that more things can go _wrong_, since you have more indexes, so \n> your verifier will have more to check, but that's not an advantage, that's \n> a downside.\n\nUh, no. It's just like a filesystem. Redundancy is what lets you\nrecover.\n\nThe extra indices are also very useful in their own right:\n\n- they let you do easily do delta storage\n- they let you efficiently do delta transmission\n- they let you find past revisions of a file in O(1)\n- they let you efficiently do \"annotate\"\n- they let you do smarter merge\n\nAt least the first four seem fairly critical to me.\n\nAs various people have pointed out, you can hack delta transmission\nand file revision indexing on top of git. But to do that, you'll need\nto build the same indices that Mercurial has. And you'll need to check\ntheir integrity.\n\nUnfortunately, since the git back-end refuses to know anything about\nthe relation between file revisions, this will all happen in the front\nend, and you'll have done almost all the work needed to do delta\nstorage without actually getting it. How sad.\n\nYou'll also likely end up with something quite a bit more complicated\nthan Mercurial because of the extra layering. This all strongly suggests\nto me that the git back-end is just a little bit too simple.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2434","messageId":"Pine.LNX.4.58.0505021932270.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"20050503000011.GA22038@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-03T02:48:29Z","receivedAt":"2005-05-03T02:48:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Matt Mackall wrote:\n> \n> Umm.. I am _not_ calculating the SHA of the delta itself. That'd be\n> silly.\n\nIt's not silly.\n\nMeta-data consistency is supremely important. If people can corrupt their \nmetadata in strange an unobservable ways, that's almost as bad as \ncorrupting the data itself. In fact, to some degree it's worse, since you \nmake people trust the thing, but you don't actually guarantee it.\n\nSo how _do_ you guarantee consistency of a tree and the history that led \nup to it? \n\nAnd by that I don't mean any of the individual blobs - I realize that it's \nperfectly valid to just check out every single version, and have the sha1 \nof that. But how do you guarantee that the sha's you check are the sha's \nthat you saved in the first place, and somebody didn't replace something \nin the middle?\n\nIn other words, you need to hash the metadata too. Otherwise how do you\nconsistency-check the _collection_ of files?\n\nIt's absolutely not enough to just protect single-file content. That \ndoesn't help one whit. It's not what a SCM is all about. You have to \nprotect the state of _multiple_ files, ie the metadata has to be \nverifiable too.\n\nIf that meta-data is the index, then the index needs to be protected by a\nSHA1. In git, that's why we don't just sha1 every blob, but every tree and\nevery commit. That's the thing that gets consistency _beyond_ a single\nfile.\n\n> As various people have pointed out, you can hack delta transmission\n> and file revision indexing on top of git. But to do that, you'll need\n> to build the same indices that Mercurial has. And you'll need to check\n> their integrity.\n\nNo, absolutely not.\n\nBuilding indeces on top of git would be stupid. You can _cache_ deltas,\nbut there's a big difference between a index that actually describes how\nrandom blobs go together, and a cache of a delta between two\nwell-specified end-points. And in particular, there is no \"consistency\" to\na delta. You don't need it.\n\nWhy? Because either the delta is correct, or it isn't. If it's correct,\nthe end result will be the right sha1. If it's not, the end result will be\nsomething else. So when you do a \"pull\" from another repository, you can\ntrivially check whether the delta's you got were valid: did applying them\nresult in the same sha1 that the other repository had?\n\nSo git really validates the _only_ thing that matters: it validates the \nstate of the data. It doesn't validate anything else, but if validates \nthat one thing very completely indeed.\n\n\t\tLinus\n"},{"id":"2439","messageId":"20050503032916.GE22038@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505021932270.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-03T03:29:16Z","receivedAt":"2005-05-03T03:29:16Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Mon, May 02, 2005 at 07:48:29PM -0700, Linus Torvalds wrote:\n> \n> \n> On Mon, 2 May 2005, Matt Mackall wrote:\n> > \n> > Umm.. I am _not_ calculating the SHA of the delta itself. That'd be\n> > silly.\n> \n> It's not silly.\n\nThe delta is not the object I care about and its representation is\narbitrary. In fact different branches will store different deltas\ndepending on how their DAGs get topologically sorted. The object I\ncare about is the original text, so that's the hash I store.\n\n> In other words, you need to hash the metadata too. Otherwise how do you\n> consistency-check the _collection_ of files?\n\nWell naturally, I hash the metadata too. For every change, there's a\ntoplevel changeset hash that is the hash of the entire project state\nat that time. And it's all signable and so on. Just like git and just\nlike Monotone.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2444","messageId":"Pine.LNX.4.58.0505022116080.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"20050503032916.GE22038@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-03T04:18:23Z","receivedAt":"2005-05-03T04:18:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Matt Mackall wrote:\n> \n> The delta is not the object I care about and its representation is\n> arbitrary. In fact different branches will store different deltas\n> depending on how their DAGs get topologically sorted. The object I\n> care about is the original text, so that's the hash I store.\n\nOk. In that case, it sounds like you're really doing everything git is\ndoing, except your \"blob\" objects effectively can have pointers to a\nprevious object (and you have a different on-disk representation)?  Is\nthat correct?\n\n\t\t\tLinus\n"},{"id":"2447","messageId":"Pine.LNX.4.58.0505022123270.3594@ppc970.osdl.org","threadId":"317","inReplyTo":"20050503000011.GA22038@waste.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-05-03T04:24:54Z","receivedAt":"2005-05-03T04:24:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 2 May 2005, Matt Mackall wrote:\n> \n> It's still simple in Mercurial, but more importantly Mercurial _won't\n> need it_. Dropping history is a work-around, not a feature.\n\nSide note: this is what Larry thought about BK too. Until three years had\npassed, and the ChangeSet file was many megabytes in size. Even slow\ngrowth ends up being big growth in the end..\n\nWe had been talking about pruning the BK history as long back as a year \nago.\n\n\t\tLinus\n"},{"id":"2449","messageId":"20050503042739.GF22038@waste.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505022123270.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-03T04:27:39Z","receivedAt":"2005-05-03T04:27:39Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Mon, May 02, 2005 at 09:24:54PM -0700, Linus Torvalds wrote:\n> \n> \n> On Mon, 2 May 2005, Matt Mackall wrote:\n> > \n> > It's still simple in Mercurial, but more importantly Mercurial _won't\n> > need it_. Dropping history is a work-around, not a feature.\n> \n> Side note: this is what Larry thought about BK too. Until three years had\n> passed, and the ChangeSet file was many megabytes in size. Even slow\n> growth ends up being big growth in the end..\n> \n> We had been talking about pruning the BK history as long back as a year \n> ago.\n\nOk, I'll implement it on my red eye flight tonight. But Mercurial\nwon't suffer from the O(filesize) problem of BK.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"2459","messageId":"20050503084543.GA26234@taniwha.stupidest.org","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505022123270.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Chris Wedgwood","fromEmail":"cw@f00f.org","sentAt":"2005-05-03T08:45:43Z","receivedAt":"2005-05-03T08:45:43Z","isPatch":false,"sender":{"key":"cw@f00f.org","avatar":null},"body":"On Mon, May 02, 2005 at 09:24:54PM -0700, Linus Torvalds wrote:\n\n> We had been talking about pruning the BK history as long back as a\n> year ago.\n\nWas that the history or all the deletes/renames that were painful\nthough?\n"},{"id":"2509","messageId":"4277B778.5020206@tmr.com","threadId":"317","inReplyTo":"200505021614.j42GEufG008441@turing-police.cc.vt.edu","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-05-03T17:40:08Z","receivedAt":"2005-05-03T17:40:08Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Valdis.Kletnieks@vt.edu wrote:\n> On Mon, 02 May 2005 11:49:32 EDT, Bill Davidsen said:\n> \n>>Andrea Arcangeli wrote:\n>>\n>>>On Fri, Apr 29, 2005 at 01:39:59PM -0700, Matt Mackall wrote:\n> \n> \n>>>-#!/usr/bin/python\n>>>+#!/usr/bin/env python\n>>> #\n>>> # mercurial - a minimal scalable distributed SCM\n>>> # v0.4b \"oedipa maas\"\n>>\n>>Could you explain why this is necessary or desirable? I looked at what \n>>env does, and I am missing the point of duplicating bash normal \n>>behaviour regarding definition of per-process environment entries.\n> \n> \n> Most likely, his python lives elsewhere than /usr/bin, and the 'env' call\n> results in causing a walk across $PATH to find it....\n\nAssuming that he has env in a standard place... I hope this isn't going \nto start some rash of efforts to make packages run on non-standard \ntoolchains, which add requirements for one tool to get around \nmisplacement of another.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n\n"},{"id":"2510","messageId":"4277B84C.8030808@tmr.com","threadId":"317","inReplyTo":"Pine.LNX.4.58.0505020921080.3594@ppc970.osdl.org","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark","fromName":"Bill Davidsen","fromEmail":"davidsen@tmr.com","sentAt":"2005-05-03T17:43:40Z","receivedAt":"2005-05-03T17:43:40Z","isPatch":false,"sender":{"key":"davidsen@tmr.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Mon, 2 May 2005, Bill Davidsen wrote:\n> \n>>>-#!/usr/bin/python\n>>>+#!/usr/bin/env python\n> \n> \n>>Could you explain why this is necessary or desirable? I looked at what \n>>env does, and I am missing the point of duplicating bash normal \n>>behaviour regarding definition of per-process environment entries.\n> \n> \n> It's not about environment.\n> \n> It's about the fact that many people have things like python in\n> /usr/local/bin/python, because they compiled it themselves or similar.\n> \n> Pretty much the only path you can _really_ depend on for #! stuff is \n> /bin/sh.\n> \n> Any system that doesn't have /bin/sh is so fucked up that it's not worth\n> worrying about. Anything else can be in /bin, /usr/bin or /usr/local/bin\n> (and sometimes other strange places).\n> \n> That said, I think the /usr/bin/env trick is stupid too. It may be more \n> portable for various Linux distributions, but if you want _true_ \n> portability, you use /bin/sh, and you do something like\n> \n> \t#!/bin/sh\n> \texec perl perlscript.pl \"$@\"\n> \n> instead.\n> \t\tLinus\n> \nAnd that eliminates the need for having /usr/bin/env in the \"expected\" \nplace. I like it.\n\nWish there was a way to specify \"use path\" without all this workaround.\n\n-- \n    -bill davidsen (davidsen@tmr.com)\n\"The secret to procrastination is to put things off until the\n  last possible moment - but no longer\"  -me\n\n"},{"id":"2555","messageId":"42782F20.5080401@dwheeler.com","threadId":"317","inReplyTo":"4277B778.5020206@tmr.com","subject":"Re: Mercurial 0.4b vs git patchbomb benchmark (/usr/bin/env again)","fromName":"David A. Wheeler","fromEmail":"dwheeler@dwheeler.com","sentAt":"2005-05-04T02:10:40Z","receivedAt":"2005-05-04T02:10:40Z","isPatch":false,"sender":{"key":"dwheeler@dwheeler.com","avatar":"https://avatars.githubusercontent.com/u/813150?v=4"},"body":"Valdis.Kletnieks@vt.edu wrote:\n>> Most likely, his python lives elsewhere than /usr/bin, and the 'env' call\n>> results in causing a walk across $PATH to find it....\n\nBill Davidsen wrote:\n> Assuming that he has env in a standard place... I hope this isn't going \n> to start some rash of efforts to make packages run on non-standard \n> toolchains, which add requirements for one tool to get around \n> misplacement of another.\n\nThe #!/usr/bin/env prefix is, in my opinion, a very good idea.\nThere are a very few systems where env isn't in /usr/bin, but they\nwere extremely rare years ago & are essentially extinct now.\nBasically, it's a 99% solution; getting the last 1% is really painful,\nbut since getting the 99% is easy, let's do it and be happy.\n\nThere are LOTS of systems where Python, bash, etc., do NOT\nlive in whatever place you think of as \"standard\".\nI routinely use an OpenBSD 3.1 system; there is no /usr/bin/bash,\nbut there _IS_ a /usr/local/bin/bash (in my PATH) and a /usr/bin/env.\nSo this /usr/bin/env stuff REALLY is useful on a lot of systems, such\nas OpenBSD.  It's critical to me, at least!\n\nThis is actually really useful on ANY system, though.\nEven if some interpreter IS where you think it should be,\nthat is NOT necessarily the interpreter you want to use.\nUsing \"/usr/bin/env\" lets you use PATH\nto override things, so you don't HAVE to use the interpreter\nin some fixed location.  That's REALLY handy for testing... I\ncan download the whizbang Python 9.8.2, set it on the path,\nand see if everything works as expected.  It's also nice\nif someone insists on never upgrading a package; you can\ninstall an interpreter \"locally\".  Yes, you can patch all the\nfiles up, but resetting a PATH is _much_ easier.\n\n--- David A. Wheeler\n"}]}