{"thread":{"id":"582","subject":"Mercurial 0.4e vs git network pull","startedAt":"2005-05-12T09:44:06Z","lastAt":"2005-05-15T08:54:05Z","messageCount":14,"participants":["Matt Mackall","Petr Baudis","Daniel Barkalow","Ingo Molnar"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"3130","messageId":"20050512094406.GZ5914@waste.org","threadId":"582","inReplyTo":null,"subject":"Mercurial 0.4e vs git network pull","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-12T09:44:06Z","receivedAt":"2005-05-12T09:44:06Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"Now that I'm back from vacation, there's a new Mercurial release as\nwell as snapshots at:\n\n  http://selenic.com/mercurial/\n\nA combined self-hosting repository / web interface can be found at:\n\n  http://selenic.com/hg/\n\nAnd there's now a mailing list at:\n\n  http://selenic.com/mailman/listinfo/mercurial\n\nThe big news is that Mercurial now has a very fast network protocol.\nThis benchmark is pulling and merging 819 changesets (again, taken\nfrom 2.6.12-rc2-mm3) from one repo to another over DSL using\nMercurial's new delta protocol:\n\n $ time hg merge hg://selenic.com/linux-hg/\n retrieving changegroup\n merging changesets\n merging manifests\n merging files\n\n real    0m10.276s\n user    0m3.299s\n sys     0m0.689s\n\nFor comparison, rsyncing the same set of changes between git repos from\nthe same server:\n\n $ time rsync -a rsync://10.0.0.12:2000/git/lgb/.git .\n sent 171508 bytes  received 31225542 bytes  312408.46 bytes/sec\n\n real    1m40.470s\n user    0m0.655s\n sys     0m1.896s\n\nThe original broken-out.tar.bz2: 2.3M\nThe same, uncompressed:           15M\nThe same, rsynced with git:       30M\nThe same, pulled with hg (zlib): 2.5M  <- what I used above\nThe same, pulled with hg (bz2):  2.1M\n\nThe server in question is a relatively busy 1GHz Athlon. The server\nside of the hg protocol is stateless and is serviced by a simple CGI\nscript run under Apache.\n\nMercurial is more than 10 times as bandwidth efficient and\nconsiderably more I/O efficient. On the server side, rsync uses about\ntwice as much CPU time as the Mercurial server and has about 10 times\nthe I/O and pagecache footprint as well.\n\nMercurial is also much smarter than rsync at determining what\noutstanding changesets exist. Here's an empty pull as a demonstration:\n\n $ time hg merge hg://selenic.com/linux-hg/\n retrieving changegroup\n\n real    0m0.363s\n user    0m0.083s\n sys     0m0.007s\n\nThat's a single http request and a one line response.\n\nAnd now with rsync:\n\n $ time rsync -av rsync://10.0.0.12:2000/git/lgb/.git .\n receiving file list ... done\n\n sent 76 bytes  received 1280245 bytes  2560642.00 bytes/sec\n total size is 85993841  speedup is 67.17\n\n real    0m0.539s\n user    0m0.185s\n sys     0m0.148s\n\nMercurial's communication here scales O(min(changed branches, log new\nchangesets)) which is less than O(new changesets), while rsync scales\nwith O(total number of file revisions) (ouch!). The above transfer\nsize for an empty pull will go from 1.2M to >12M when there's similar\nhistory in git to what's in BK.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"3172","messageId":"20050512182340.GA324@pasky.ji.cz","threadId":"582","inReplyTo":"20050512094406.GZ5914@waste.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-05-12T18:23:41Z","receivedAt":"2005-05-12T18:23:41Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, May 12, 2005 at 11:44:06AM CEST, I got a letter\nwhere Matt Mackall <mpm@selenic.com> told me that...\n> Mercurial is more than 10 times as bandwidth efficient and\n> considerably more I/O efficient. On the server side, rsync uses about\n> twice as much CPU time as the Mercurial server and has about 10 times\n> the I/O and pagecache footprint as well.\n> \n> Mercurial is also much smarter than rsync at determining what\n> outstanding changesets exist. Here's an empty pull as a demonstration:\n> \n>  $ time hg merge hg://selenic.com/linux-hg/\n>  retrieving changegroup\n> \n>  real    0m0.363s\n>  user    0m0.083s\n>  sys     0m0.007s\n> \n> That's a single http request and a one line response.\n\nSo, what about comparing it with something comparable, say git pull over\nHTTP? :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"3187","messageId":"20050512201116.GC5914@waste.org","threadId":"582","inReplyTo":"20050512182340.GA324@pasky.ji.cz","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-12T20:11:16Z","receivedAt":"2005-05-12T20:11:16Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Thu, May 12, 2005 at 08:23:41PM +0200, Petr Baudis wrote:\n> Dear diary, on Thu, May 12, 2005 at 11:44:06AM CEST, I got a letter\n> where Matt Mackall <mpm@selenic.com> told me that...\n> > Mercurial is more than 10 times as bandwidth efficient and\n> > considerably more I/O efficient. On the server side, rsync uses about\n> > twice as much CPU time as the Mercurial server and has about 10 times\n> > the I/O and pagecache footprint as well.\n> > \n> > Mercurial is also much smarter than rsync at determining what\n> > outstanding changesets exist. Here's an empty pull as a demonstration:\n> > \n> >  $ time hg merge hg://selenic.com/linux-hg/\n> >  retrieving changegroup\n> > \n> >  real    0m0.363s\n> >  user    0m0.083s\n> >  sys     0m0.007s\n> > \n> > That's a single http request and a one line response.\n> \n> So, what about comparing it with something comparable, say git pull over\n> HTTP? :-)\n\n..because I get a headache every time I try to figure out how to use git? :-P\n\nSeriously, have a pointer to how this works?\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"3189","messageId":"20050512201406.GJ324@pasky.ji.cz","threadId":"582","inReplyTo":"20050512201116.GC5914@waste.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-05-12T20:14:06Z","receivedAt":"2005-05-12T20:14:06Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, May 12, 2005 at 10:11:16PM CEST, I got a letter\nwhere Matt Mackall <mpm@selenic.com> told me that...\n> On Thu, May 12, 2005 at 08:23:41PM +0200, Petr Baudis wrote:\n> > Dear diary, on Thu, May 12, 2005 at 11:44:06AM CEST, I got a letter\n> > where Matt Mackall <mpm@selenic.com> told me that...\n> > > Mercurial is more than 10 times as bandwidth efficient and\n> > > considerably more I/O efficient. On the server side, rsync uses about\n> > > twice as much CPU time as the Mercurial server and has about 10 times\n> > > the I/O and pagecache footprint as well.\n> > > \n> > > Mercurial is also much smarter than rsync at determining what\n> > > outstanding changesets exist. Here's an empty pull as a demonstration:\n> > > \n> > >  $ time hg merge hg://selenic.com/linux-hg/\n> > >  retrieving changegroup\n> > > \n> > >  real    0m0.363s\n> > >  user    0m0.083s\n> > >  sys     0m0.007s\n> > > \n> > > That's a single http request and a one line response.\n> > \n> > So, what about comparing it with something comparable, say git pull over\n> > HTTP? :-)\n> \n> ..because I get a headache every time I try to figure out how to use git? :-P\n> \n> Seriously, have a pointer to how this works?\n\nEither you use cogito and just pass cg-clone an HTTP URL (to the git\nrepository as in the case of rsync -\nhttp://www.kernel.org/pub/scm/cogito/cogito.git should work), or you\ninvoke git-http-pull directly (passing it desired commit ID of the\nremote HEAD you want to fetch, and the URL; see\nDocumentation/git-http-pull.txt).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"3196","messageId":"20050512205735.GE5914@waste.org","threadId":"582","inReplyTo":"20050512201406.GJ324@pasky.ji.cz","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-12T20:57:35Z","receivedAt":"2005-05-12T20:57:35Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Thu, May 12, 2005 at 10:14:06PM +0200, Petr Baudis wrote:\n> Dear diary, on Thu, May 12, 2005 at 10:11:16PM CEST, I got a letter\n> where Matt Mackall <mpm@selenic.com> told me that...\n> > On Thu, May 12, 2005 at 08:23:41PM +0200, Petr Baudis wrote:\n> > > Dear diary, on Thu, May 12, 2005 at 11:44:06AM CEST, I got a letter\n> > > where Matt Mackall <mpm@selenic.com> told me that...\n> > > > Mercurial is more than 10 times as bandwidth efficient and\n> > > > considerably more I/O efficient. On the server side, rsync uses about\n> > > > twice as much CPU time as the Mercurial server and has about 10 times\n> > > > the I/O and pagecache footprint as well.\n> > > > \n> > > > Mercurial is also much smarter than rsync at determining what\n> > > > outstanding changesets exist. Here's an empty pull as a demonstration:\n> > > > \n> > > >  $ time hg merge hg://selenic.com/linux-hg/\n> > > >  retrieving changegroup\n> > > > \n> > > >  real    0m0.363s\n> > > >  user    0m0.083s\n> > > >  sys     0m0.007s\n> > > > \n> > > > That's a single http request and a one line response.\n> > > \n> > > So, what about comparing it with something comparable, say git pull over\n> > > HTTP? :-)\n> > \n> > ..because I get a headache every time I try to figure out how to use git? :-P\n> > \n> > Seriously, have a pointer to how this works?\n> \n> Either you use cogito and just pass cg-clone an HTTP URL (to the git\n> repository as in the case of rsync -\n> http://www.kernel.org/pub/scm/cogito/cogito.git should work), or you\n> invoke git-http-pull directly (passing it desired commit ID of the\n> remote HEAD you want to fetch, and the URL; see\n> Documentation/git-http-pull.txt).\n\nDoes this need an HTTP request (and round trip) per object? It appears\nto. That's 2200 requests/round trips for my 800 patch benchmark.\n\nHow does git find the outstanding changesets?\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"3202","messageId":"Pine.LNX.4.21.0505121709250.30848-100000@iabervon.org","threadId":"582","inReplyTo":"20050512205735.GE5914@waste.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-05-12T21:24:27Z","receivedAt":"2005-05-12T21:24:27Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 12 May 2005, Matt Mackall wrote:\n\n> Does this need an HTTP request (and round trip) per object? It appears\n> to. That's 2200 requests/round trips for my 800 patch benchmark.\n\nIt requires a request per object, but it should be possible (with\nsomewhat more complicated code) to overlap them such that it doesn't\nrequire a serial round trip for each. Since the server is sending static\nfiles, the overhead for each should be minimal.\n\n> How does git find the outstanding changesets?\n\nIn the present mainline, you first have to find the head commit you\nwant. I have a patch which does this for you over the same\nconnection. Starting from that point, it tracks reachability on the\nreceiving end, and requests anything it doesn't have.\n\nFor the case of having nothing to do, it should be a single one-line\nrequest/response for a static file (after which the local end determines\nthat it has everything it needs without talking to the server).\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"3211","messageId":"20050512222943.GI5914@waste.org","threadId":"582","inReplyTo":"Pine.LNX.4.21.0505121709250.30848-100000@iabervon.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-12T22:29:43Z","receivedAt":"2005-05-12T22:29:43Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Thu, May 12, 2005 at 05:24:27PM -0400, Daniel Barkalow wrote:\n> On Thu, 12 May 2005, Matt Mackall wrote:\n> \n> > Does this need an HTTP request (and round trip) per object? It appears\n> > to. That's 2200 requests/round trips for my 800 patch benchmark.\n> \n> It requires a request per object, but it should be possible (with\n> somewhat more complicated code) to overlap them such that it doesn't\n> require a serial round trip for each. Since the server is sending static\n> files, the overhead for each should be minimal.\n\nIt's not minimal. The size of an HTTP request is often not much\ndifferent than the size of a compressed file delta. Here's one of the\nindexes from a file in an hg repo:\n\n   rev    offset  length  base linkrev p1           p2           nodeid\n     0         0    2307     0       0 0000000000.. 0000000000.. b6444347c6..\n     1      2307      77     0       5 b6444347c6.. 0000000000.. 06763db6de..\n     2      2384     225     0      11 06763db6de.. 0000000000.. acc8e2b2f0..\n     3      2609      40     0      16 acc8e2b2f0.. 0000000000.. 461b079d98..\n     4      2649     261     0      17 461b079d98.. 0000000000.. 8507ba44cc..\n     5      2910     486     0      18 8507ba44cc.. 0000000000.. b68523252b..\n     6      3396      98     0      21 b68523252b.. 0000000000.. b3f2586243..\n     7      3494     238     0      22 b3f2586243.. 0000000000.. d73d0f8ee9..\n     8      3732      39     0      23 d73d0f8ee9.. 0000000000.. caaf506196..\n     9      3771     266     0      24 caaf506196.. 0000000000.. 54485fc96f..\n    10      4037      81     0      29 54485fc96f.. 0000000000.. b9eae7b990..\n    11      4118     310     0      31 b9eae7b990.. 0000000000.. a9926b092a..\n    12      4428     545     0      33 a9926b092a.. 0000000000.. f26c600172..\n    13      4973     419     0      34 f26c600172.. 0000000000.. ec4ab0acb7..\n    14      5392     136     0      38 ec4ab0acb7.. 0000000000.. eb5f3f76c8..\n    15      5528     161     0      39 eb5f3f76c8.. 0000000000.. 4fc5f3a3ae..\n    16      5689     258     0      46 4fc5f3a3ae.. 0000000000.. 3ad83891fb..\n    17      5947     171     0      49 3ad83891fb.. 0000000000.. 3983ac6cd2..\n    18      6118     195     0      50 3983ac6cd2.. 0000000000.. f138865e04..\n    19      6313      79     0      52 f138865e04.. 0000000000.. 3566c1f449..\n    20      6392      85     0      53 3566c1f449.. 0000000000.. 0694a4e3eb..\n    21      6477      91     0      54 0694a4e3eb.. 0000000000.. 5f98ae7426..\n    22      6568     208     0      56 5f98ae7426.. 0000000000.. dae5cb80db..\n    23      6776     286     0      62 dae5cb80db.. 0000000000.. 90ff243869..\n\nAll the junk that gets bundled in an http request/response will be\nsimilar in size to the stuff in the third column.\n\nRelative to the 10-20x overhead of not sending deltas, yes, it's only 10%.\n \n> > How does git find the outstanding changesets?\n> \n> In the present mainline, you first have to find the head commit you\n> want. I have a patch which does this for you over the same\n> connection. Starting from that point, it tracks reachability on the\n> receiving end, and requests anything it doesn't have.\n\nDoes it do this recursively? Eg, if the server has 800 new linear\ncommits, does the client have to do 800 round trips following parent\npointers to find all the new changesets? In this case, Mercurial does\nabout 6 round trips, totalling less than 1K, plus one requests\nthat pulls everything.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"3220","messageId":"Pine.LNX.4.21.0505121949210.30848-100000@iabervon.org","threadId":"582","inReplyTo":"20050512222943.GI5914@waste.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-05-13T00:33:56Z","receivedAt":"2005-05-13T00:33:56Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 12 May 2005, Matt Mackall wrote:\n\n> On Thu, May 12, 2005 at 05:24:27PM -0400, Daniel Barkalow wrote:\n> > On Thu, 12 May 2005, Matt Mackall wrote:\n> > \n> > > Does this need an HTTP request (and round trip) per object? It appears\n> > > to. That's 2200 requests/round trips for my 800 patch benchmark.\n> > \n> > It requires a request per object, but it should be possible (with\n> > somewhat more complicated code) to overlap them such that it doesn't\n> > require a serial round trip for each. Since the server is sending static\n> > files, the overhead for each should be minimal.\n> \n> It's not minimal. The size of an HTTP request is often not much\n> different than the size of a compressed file delta.\n\nI was thinking of server-side processing overhead, not bandwidth. It's\ntrue that the bandwidth could be noticeable for these small files.\n\n> All the junk that gets bundled in an http request/response will be\n> similar in size to the stuff in the third column.\n\nkernel.org seems to send 283-byte responses, to be completely\nprecise. This could be cut down substantially if Apache were tweaked a bit\nto skip all the optional headers which are useless or wrong in this\ncontext. (E.g., that includes sending a content-type of \"text/plain\" for\nthe binary data)\n\n> Does it do this recursively? Eg, if the server has 800 new linear\n> commits, does the client have to do 800 round trips following parent\n> pointers to find all the new changesets? \n\nYes, although that also includes pulling the commits, and may be\ninterleaved with pulling the trees and objects to cover the\nlatency. (I.e., one round trip gets the new head hash; the second gets\nthat commit; on the third the tree and the parent(s) can be requested at\nonce; on the fouth the contents of the tree and the grandparents, at\nwhich point the bandwidth will probably be the limiting factor for the\nrest of the operation.)\n\n> In this case, Mercurial does about 6 round trips, totalling less than\n> 1K, plus one requests that pulls everything.\n\nI must be misunderstanding your numbers, because 6 HTTP responses is more\nthan 1K, ignoring any actual content from the server, and 1K for 800\ncommits is less than 2 bytes per commit.\n\nI'm also worried about testing on 800 linear commits, since the projects\nunder consideration tend to have very non-linear histories. \n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"3221","messageId":"20050513011149.GK5914@waste.org","threadId":"582","inReplyTo":"Pine.LNX.4.21.0505121949210.30848-100000@iabervon.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-13T01:11:49Z","receivedAt":"2005-05-13T01:11:49Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Thu, May 12, 2005 at 08:33:56PM -0400, Daniel Barkalow wrote:\n> On Thu, 12 May 2005, Matt Mackall wrote:\n> \n> > On Thu, May 12, 2005 at 05:24:27PM -0400, Daniel Barkalow wrote:\n> > > On Thu, 12 May 2005, Matt Mackall wrote:\n> > > \n> > > > Does this need an HTTP request (and round trip) per object? It appears\n> > > > to. That's 2200 requests/round trips for my 800 patch benchmark.\n> > > \n> > > It requires a request per object, but it should be possible (with\n> > > somewhat more complicated code) to overlap them such that it doesn't\n> > > require a serial round trip for each. Since the server is sending static\n> > > files, the overhead for each should be minimal.\n> > \n> > It's not minimal. The size of an HTTP request is often not much\n> > different than the size of a compressed file delta.\n> \n> I was thinking of server-side processing overhead, not bandwidth. It's\n> true that the bandwidth could be noticeable for these small files.\n> \n> > All the junk that gets bundled in an http request/response will be\n> > similar in size to the stuff in the third column.\n> \n> kernel.org seems to send 283-byte responses, to be completely\n> precise. This could be cut down substantially if Apache were tweaked a bit\n> to skip all the optional headers which are useless or wrong in this\n> context. (E.g., that includes sending a content-type of \"text/plain\" for\n> the binary data)\n> \n> > Does it do this recursively? Eg, if the server has 800 new linear\n> > commits, does the client have to do 800 round trips following parent\n> > pointers to find all the new changesets? \n> \n> Yes, although that also includes pulling the commits, and may be\n> interleaved with pulling the trees and objects to cover the\n> latency. (I.e., one round trip gets the new head hash; the second gets\n> that commit; on the third the tree and the parent(s) can be requested at\n> once; on the fouth the contents of the tree and the grandparents, at\n> which point the bandwidth will probably be the limiting factor for the\n> rest of the operation.)\n\nWhat if a changeset is smaller than the bandwidth-delay product of\nyour link? As an extreme example, Mercurial is currently at a point\nwhere its -entire repo- changegroup (set of all changesets) can be in\nflight on the wire on a typical link.\n\n> > In this case, Mercurial does about 6 round trips, totalling less than\n> > 1K, plus one requests that pulls everything.\n> \n> I must be misunderstanding your numbers, because 6 HTTP responses is more\n> than 1K, ignoring any actual content from the server, and 1K for 800\n> commits is less than 2 bytes per commit.\n\n1k of application-level data, sorry. And my whole point is that I\ndon't send those 800 commit identifiers (which are 40 bytes each as\nhex). I send about 30 or so. It's basically a negotiation to find the\nearliest commits not known to the client with a minimum of round trips\nand data exchange.\n\n> I'm also worried about testing on 800 linear commits, since the projects\n> under consideration tend to have very non-linear histories. \n\nNot true at all. Dumps from Andrew to Linus via patch bombs will\nresult in runs of hundreds of linear commits on a regular basis.\nLinear patch series are the preferred way to make changes and series\nof 30 or 40 small patches are not at all uncommon.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"3223","messageId":"Pine.LNX.4.21.0505122148480.30848-100000@iabervon.org","threadId":"582","inReplyTo":"20050513011149.GK5914@waste.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-05-13T02:23:01Z","receivedAt":"2005-05-13T02:23:01Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 12 May 2005, Matt Mackall wrote:\n\n> On Thu, May 12, 2005 at 08:33:56PM -0400, Daniel Barkalow wrote:\n>\n> > Yes, although that also includes pulling the commits, and may be\n> > interleaved with pulling the trees and objects to cover the\n> > latency. (I.e., one round trip gets the new head hash; the second gets\n> > that commit; on the third the tree and the parent(s) can be requested at\n> > once; on the fouth the contents of the tree and the grandparents, at\n> > which point the bandwidth will probably be the limiting factor for the\n> > rest of the operation.)\n> \n> What if a changeset is smaller than the bandwidth-delay product of\n> your link? As an extreme example, Mercurial is currently at a point\n> where its -entire repo- changegroup (set of all changesets) can be in\n> flight on the wire on a typical link.\n\nIf this is common for the repository in question, then it will be forced\nto wait for the parent to come in, true. If you have a number of merges,\nhowever, you start using more total bandwidth relative to latency while\ntracking them in parallel.\n\n> > I must be misunderstanding your numbers, because 6 HTTP responses is more\n> > than 1K, ignoring any actual content from the server, and 1K for 800\n> > commits is less than 2 bytes per commit.\n> \n> 1k of application-level data, sorry. And my whole point is that I\n> don't send those 800 commit identifiers (which are 40 bytes each as\n> hex). I send about 30 or so. It's basically a negotiation to find the\n> earliest commits not known to the client with a minimum of round trips\n> and data exchange.\n\nDoes this rely on the history being entirely linear? I suppose that\nrequesting a rev-list from the server (which could have it as a static\nfile generated when a new head was pushed) could jumpstart the\nprocess. The client could request all of the commits it doesn't have in\nrapid succession, and then request trees as the commits started coming\nin. Of course, this would get inefficient if you were, for example,\npulling a merge with a branch with a long history, since you'd get a ton\nof old mainline (which you already have) interleaved with occasional new\nthings.\n\n> > I'm also worried about testing on 800 linear commits, since the projects\n> > under consideration tend to have very non-linear histories. \n> \n> Not true at all. Dumps from Andrew to Linus via patch bombs will\n> result in runs of hundreds of linear commits on a regular basis.\n> Linear patch series are the preferred way to make changes and series\n> of 30 or 40 small patches are not at all uncommon.\n\nIt has sounded like Andrew had some interest in using git, and a number of\nother developers are using it already. If this becomes still more common,\nit may be the case that, instead of sending patch bombs, Andrew will point\nLinus at authors' original series, in which case the mainline would be\nmerges of a hundred linear series of various lengths. I had the\nimpression, although I never looked carefully, that this was happening on\na smaller scale with BK, where work by BK users got included using BK,\nrather than as patches applied out of a bomb.\n\nIt certainly makes sense as a design goal to be able to support everything\nhappening within the system, rather than getting exported and reimported.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n"},{"id":"3224","messageId":"20050513024427.GL5914@waste.org","threadId":"582","inReplyTo":"Pine.LNX.4.21.0505122148480.30848-100000@iabervon.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Matt Mackall","fromEmail":"mpm@selenic.com","sentAt":"2005-05-13T02:44:27Z","receivedAt":"2005-05-13T02:44:27Z","isPatch":false,"sender":{"key":"mpm@selenic.com","avatar":null},"body":"On Thu, May 12, 2005 at 10:23:01PM -0400, Daniel Barkalow wrote:\n> On Thu, 12 May 2005, Matt Mackall wrote:\n> \n> > On Thu, May 12, 2005 at 08:33:56PM -0400, Daniel Barkalow wrote:\n> >\n> > > Yes, although that also includes pulling the commits, and may be\n> > > interleaved with pulling the trees and objects to cover the\n> > > latency. (I.e., one round trip gets the new head hash; the second gets\n> > > that commit; on the third the tree and the parent(s) can be requested at\n> > > once; on the fouth the contents of the tree and the grandparents, at\n> > > which point the bandwidth will probably be the limiting factor for the\n> > > rest of the operation.)\n> > \n> > What if a changeset is smaller than the bandwidth-delay product of\n> > your link? As an extreme example, Mercurial is currently at a point\n> > where its -entire repo- changegroup (set of all changesets) can be in\n> > flight on the wire on a typical link.\n> \n> If this is common for the repository in question, then it will be forced\n> to wait for the parent to come in, true. If you have a number of merges,\n> however, you start using more total bandwidth relative to latency while\n> tracking them in parallel.\n\nNo, you're missing my point. If you can request all the files in a\nchangeset in less than a round-trip time, you have a pipeline stall.\nLet's say a changeset is 10k and round trip time is 100ms. That means\nyou'll stall on any pipe with more than 100k/s. You won't know what\nchangeset to request next as it'll still be in flight.\n \n> > > I must be misunderstanding your numbers, because 6 HTTP responses is more\n> > > than 1K, ignoring any actual content from the server, and 1K for 800\n> > > commits is less than 2 bytes per commit.\n> > \n> > 1k of application-level data, sorry. And my whole point is that I\n> > don't send those 800 commit identifiers (which are 40 bytes each as\n> > hex). I send about 30 or so. It's basically a negotiation to find the\n> > earliest commits not known to the client with a minimum of round trips\n> > and data exchange.\n> \n> Does this rely on the history being entirely linear? I suppose that\n> requesting a rev-list from the server (which could have it as a static\n> file generated when a new head was pushed) could jumpstart the\n> process. The client could request all of the commits it doesn't have in\n> rapid succession, and then request trees as the commits started coming\n> in. Of course, this would get inefficient if you were, for example,\n> pulling a merge with a branch with a long history, since you'd get a ton\n> of old mainline (which you already have) interleaved with occasional new\n> things.\n\nI don't depend on history being linear (I'm not reinventing CVS here)\nand I don't grab a list of all revisions (the point is to be\nscalable). In fact, I do something fairly clever, and something I\ndon't think will work with git, because, yet again, it lacks the\nmetadata.\n\n> > > I'm also worried about testing on 800 linear commits, since the projects\n> > > under consideration tend to have very non-linear histories. \n> > \n> > Not true at all. Dumps from Andrew to Linus via patch bombs will\n> > result in runs of hundreds of linear commits on a regular basis.\n> > Linear patch series are the preferred way to make changes and series\n> > of 30 or 40 small patches are not at all uncommon.\n> \n> It has sounded like Andrew had some interest in using git, and a number of\n> other developers are using it already. If this becomes still more common,\n> it may be the case that, instead of sending patch bombs, Andrew will point\n> Linus at authors' original series, in which case the mainline would be\n> merges of a hundred linear series of various lengths. I had the\n> impression, although I never looked carefully, that this was happening on\n> a smaller scale with BK, where work by BK users got included using BK,\n> rather than as patches applied out of a bomb.\n\nAndrew already uses git, in a manner much like he used BK. He does a\npull from a repo, generates a patch of that repo vs mainline, and puts\nthat in -mm. And never passes that stuff on to Linus.\n\n-- \nMathematics is the supreme nostalgia of our time.\n"},{"id":"3229","messageId":"20050513054424.GG16464@pasky.ji.cz","threadId":"582","inReplyTo":"Pine.LNX.4.21.0505121709250.30848-100000@iabervon.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-05-13T05:44:24Z","receivedAt":"2005-05-13T05:44:24Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, May 12, 2005 at 11:24:27PM CEST, I got a letter\nwhere Daniel Barkalow <barkalow@iabervon.org> told me that...\n> In the present mainline, you first have to find the head commit you\n> want. I have a patch which does this for you over the same\n> connection. Starting from that point, it tracks reachability on the\n> receiving end, and requests anything it doesn't have.\n\nCould we get the patch, please? :-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"},{"id":"3343","messageId":"20050515062243.GA22021@elte.hu","threadId":"582","inReplyTo":"20050512182340.GA324@pasky.ji.cz","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Ingo Molnar","fromEmail":"mingo@elte.hu","sentAt":"2005-05-15T06:22:43Z","receivedAt":"2005-05-15T06:22:43Z","isPatch":false,"sender":{"key":"mingo@elte.hu","avatar":null},"body":"\n* Petr Baudis <pasky@ucw.cz> wrote:\n\n> > Mercurial is also much smarter than rsync at determining what\n> > outstanding changesets exist. Here's an empty pull as a demonstration:\n> > \n> >  $ time hg merge hg://selenic.com/linux-hg/\n> >  retrieving changegroup\n> > \n> >  real    0m0.363s\n> >  user    0m0.083s\n> >  sys     0m0.007s\n> > \n> > That's a single http request and a one line response.\n> \n> So, what about comparing it with something comparable, say git pull \n> over HTTP? :-)\n\nMatt, did you get around to do such a comparison?\n\n\tIngo\n"},{"id":"3349","messageId":"20050515085405.GB13024@pasky.ji.cz","threadId":"582","inReplyTo":"20050512205735.GE5914@waste.org","subject":"Re: Mercurial 0.4e vs git network pull","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-05-15T08:54:05Z","receivedAt":"2005-05-15T08:54:05Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, May 12, 2005 at 10:57:35PM CEST, I got a letter\nwhere Matt Mackall <mpm@selenic.com> told me that...\n> Does this need an HTTP request (and round trip) per object? It appears\n> to. That's 2200 requests/round trips for my 800 patch benchmark.\n\nYes it does. On the other side, it needs no server-side CGI. But I guess\nit should be pretty easy to write some kind of server-side CGI streamer,\nand it would then easily take just a single HTTP request (telling the\nserver the commit ID and receiving back all the objects).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nC++: an octopus made by nailing extra legs onto a dog. -- Steve Taylor\n"}]}