{"thread":{"id":"4753","subject":"git-fetch per-repository speed issues","startedAt":"2006-07-03T18:02:44Z","lastAt":"2006-07-06T23:36:35Z","messageCount":30,"participants":["Keith Packard","Linus Torvalds","Jeff King","Ryan Anderson","Junio C Hamano","Jakub Narebski","Andreas Ericsson","Matthias Kestenholz","Thomas Glanzmann","David Woodhouse"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"23151","messageId":"1151949764.4723.51.camel@neko.keithp.com","threadId":"4753","inReplyTo":null,"subject":"git-fetch per-repository speed issues","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-07-03T18:02:44Z","receivedAt":"2006-07-03T18:02:44Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"Ok, so maybe X.org is using git in an unexpected (or even wrong)\nfashion. Our environment has split development across dozens of separate\nrepositories which match ABI interfaces. With CVS, we were able to keep\nthis all in one giant CVS repository with separate modules, but git\ndoesn't have that notion (which is mostly good). As such, we could use\ncvsup or rsync to update the entire collection of modules.\n\nWith git, we'd prefer to use the git protocol instead of rsync for the\nusual pack-related reasons, but that is limited to a single repository\nat a time. And, it's painfully slow, even when the repository is up to\ndate:\n\n$ cd lib/libXrandr\n$ time git-fetch origin\n...\n\nreal    0m17.035s\nuser    0m2.584s\nsys     0m0.576s\n\nThis is a repository with 24 files and perhaps 50 revisions. Given\nX.org's 307 git repositories, I'll clearly need to find a faster way\nthan running git-fetch on every one.\n\nOne thing I noticed was that the git+ssh syntax found in remotes files\ndoesn't do what I thought it did -- I assumed this would use 'git' for\nfetch and 'ssh' for push, when in fact it just uses ssh for everything.\nThis slows down the connection process by several seconds.\n\n-- \nkeith.packard@intel.com\n"},{"id":"23153","messageId":"Pine.LNX.4.64.0607031603290.12404@g5.osdl.org","threadId":"4753","inReplyTo":"1151949764.4723.51.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-03T23:14:10Z","receivedAt":"2006-07-03T23:14:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Keith Packard wrote:\n> \n> With git, we'd prefer to use the git protocol instead of rsync for the\n> usual pack-related reasons, but that is limited to a single repository\n> at a time.\n\nWell, you could use multiple branches in the same repository, even if they \nare totally unrealated. That would allow you to fetch them all in one go.\n\nOne way to do that is to just name the branches hierarcially have one \nrepo, but then call the branches something like\n\n\tlibXrandr/master\n\tlibXrandr/develop\n\tXorg/master\n\tXorg/develop\n\t..\n\n> And, it's painfully slow, even when the repository is up to\n> date:\n> \n> $ cd lib/libXrandr\n> $ time git-fetch origin\n> ...\n> \n> real    0m17.035s\n> user    0m2.584s\n> sys     0m0.576s\n\nThat's _seriously_ wrong. If everything is up-to-date, a fetch should be \nbasically zero-cost. That's especially true with the anonymous git \nprotocol, which doesn't have any connection validation overhead (for the \nssh protocol, the cost is usually the ssh login).\n\nBut there may well be some bug there.\n\nLook at this:\n\n\t[torvalds@g5 git]$ time git fetch git://git.kernel.org/pub/scm/git/git.git \n\t\n\treal    0m0.431s\n\tuser    0m0.036s\n\tsys     0m0.024s\n\nand that's over my DSL line, not some studly network thing. \n\nBasically, a repo that is up-to-date should do a \"git fetch\" about as \nquickly as it does a \"git ls-remote\". Which in turn really shouldn't be \ndoing much anything at all, apart from the connect itself:\n\n\t[torvalds@g5 git]$ time git ls-remote master.kernel.org:/pub/scm/git/git.git > /dev/null \n\t\n\treal    0m1.758s\n\tuser    0m0.188s\n\tsys     0m0.024s\n\t[torvalds@g5 git]$ time git ls-remote git://git.kernel.org/pub/scm/git/git.git > /dev/null \n\t\n\treal    0m0.431s\n\tuser    0m0.056s\n\tsys     0m0.016s\n\n(note how the ssh connection is much slower - it actually ends up doing \nall the ssh back-and-forth).\n\nCan you try from different hosts? One problem may be the remote end \njust trying to do reverse DNS lookups for xinetd or whatever?\n\nAlso, one thing to try is to just do\n\n\tstrace -Ttt git-peek-remote ...\n\nwhich shows where the time is going (I selected \"git-peek-remote\", because \nthat's a simple program).\n\n\t\tLinus\n"},{"id":"23155","messageId":"20060704002138.GB5716@coredump.intra.peff.net","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607031603290.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-07-04T00:21:38Z","receivedAt":"2006-07-04T00:21:38Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 03, 2006 at 04:14:10PM -0700, Linus Torvalds wrote:\n\n> Well, you could use multiple branches in the same repository, even if they \n> are totally unrealated. That would allow you to fetch them all in one go.\n\nOne annoying thing about this is that you may want to have several of\nthe branches checked out at a time (i.e., you want the actual directory\nstructure of libXrandr/, Xorg/, etc). You could pull everything down\ninto one repo and point small pseudo-repos at it with alternates, but I\nwould think that would become a mess with pushes. You can do some magic\nwith read-tree --prefix, but again, I'm not sure how you'd make commits\non the correct branch.  Is there an easier way to do this?\n\n> Basically, a repo that is up-to-date should do a \"git fetch\" about as \n> quickly as it does a \"git ls-remote\". Which in turn really shouldn't be \n> doing much anything at all, apart from the connect itself:\n\nFetching by ssh actually makes two ssh connections (the second is to\ngrab tags).\n\n-Peff\n"},{"id":"23159","messageId":"44A9C2D2.6010409@michonline.com","threadId":"4753","inReplyTo":"20060704002138.GB5716@coredump.intra.peff.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2006-07-04T01:22:26Z","receivedAt":"2006-07-04T01:22:26Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"Jeff King wrote:\n> On Mon, Jul 03, 2006 at 04:14:10PM -0700, Linus Torvalds wrote:\n>\n>   \n>> Well, you could use multiple branches in the same repository, even if they \n>> are totally unrealated. That would allow you to fetch them all in one go.\n>>     \n>\n> One annoying thing about this is that you may want to have several of\n> the branches checked out at a time (i.e., you want the actual directory\n> structure of libXrandr/, Xorg/, etc). You could pull everything down\n> into one repo and point small pseudo-repos at it with alternates, but I\n> would think that would become a mess with pushes. You can do some magic\n> with read-tree --prefix, but again, I'm not sure how you'd make commits\n> on the correct branch.  Is there an easier way to do this?\n>   \nYou can have multiple source trees, one per 'branch' (which is a bit of\na bad term here), and have completely unrelated things in the branches.\n\nSee, for an example, the main Git repo, which has the \"man\", \"html\", and\n\"todo\" branches, logically distinct and (somewhat) unrelated to the main\nbranch tucked away in \"master\".\n\n-- \n\nRyan Anderson\n  sometimes Pug Majere\n\n\n"},{"id":"23164","messageId":"20060704014441.GB9061@coredump.intra.peff.net","threadId":"4753","inReplyTo":"44A9C2D2.6010409@michonline.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-07-04T01:44:41Z","receivedAt":"2006-07-04T01:44:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 03, 2006 at 06:22:26PM -0700, Ryan Anderson wrote:\n\n> You can have multiple source trees, one per 'branch' (which is a bit of\n> a bad term here), and have completely unrelated things in the branches.\n> \n> See, for an example, the main Git repo, which has the \"man\", \"html\", and\n> \"todo\" branches, logically distinct and (somewhat) unrelated to the main\n> branch tucked away in \"master\".\n\nRight, I know, but my complaint is that I can't then turn that into a\ndirectory hierarchy of .../man, .../html, .../todo that are all checked\nout at the same time (there are obviously ways of playing with it, say\nby setting GIT_DIR and doing a checkout in those directories, but then I\ncan't use git in the normal way).\n\nThe best I can come up with is having man, html, and todo repos pointing\nto the one (now local) repo which contains everything. But then pushing\nis a two-step process.\n\n-Peff\n"},{"id":"23168","messageId":"44A9CAA5.4050602@michonline.com","threadId":"4753","inReplyTo":"20060704014441.GB9061@coredump.intra.peff.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2006-07-04T01:55:49Z","receivedAt":"2006-07-04T01:55:49Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"Jeff King wrote:\n> On Mon, Jul 03, 2006 at 06:22:26PM -0700, Ryan Anderson wrote:\n>\n>   \n>> You can have multiple source trees, one per 'branch' (which is a bit of\n>> a bad term here), and have completely unrelated things in the branches.\n>>\n>> See, for an example, the main Git repo, which has the \"man\", \"html\", and\n>> \"todo\" branches, logically distinct and (somewhat) unrelated to the main\n>> branch tucked away in \"master\".\n>>     \n>\n> Right, I know, but my complaint is that I can't then turn that into a\n> directory hierarchy of .../man, .../html, .../todo that are all checked\n> out at the same time (there are obviously ways of playing with it, say\n> by setting GIT_DIR and doing a checkout in those directories, but then I\n> can't use git in the normal way).\n>\n> The best I can come up with is having man, html, and todo repos pointing\n> to the one (now local) repo which contains everything. But then pushing\n> is a two-step process.\n>\n>   \nHrm, if I understand CVS at all, the old workflow was \"cvsup a copy of\nthe repository, update a working tree against that\", which is, I think,\nactually even worse than the Git equivalent, since you can't reliably\neven commit to that local clone of the CVS repository.\n\nWhat am I missing?\n\nYou can still push directly upstream, I suppose, and just do 2-stage\npulls down.\n\n-- \n\nRyan Anderson\n  sometimes Pug Majere\n\n\n"},{"id":"23169","messageId":"Pine.LNX.4.64.0607032007290.12404@g5.osdl.org","threadId":"4753","inReplyTo":"20060704002138.GB5716@coredump.intra.peff.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T03:07:49Z","receivedAt":"2006-07-04T03:07:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Jeff King wrote:\n> \n> Fetching by ssh actually makes two ssh connections (the second is to\n> grab tags).\n\nTrue. Although that should happen only if there are any new tags.\n\n\t\tLinus\n"},{"id":"23170","messageId":"Pine.LNX.4.64.0607032008590.12404@g5.osdl.org","threadId":"4753","inReplyTo":"1151973438.4723.70.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T03:21:30Z","receivedAt":"2006-07-04T03:21:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Keith Packard wrote:\n> On Mon, 2006-07-03 at 16:14 -0700, Linus Torvalds wrote:\n> > \n> > Well, you could use multiple branches in the same repository, even if they \n> > are totally unrealated. That would allow you to fetch them all in one go.\n> \n> I'd like to avoid this; the hope is that most people won't ever need to\n> look at most repositories; it would be somewhat like having glibc in the\n> same repo as the kernel...\n\nSure, understood. I'm just saying that if you want to fetch in one go, \nit's one possibility.\n\nHowever, your setup has something else seriously wrong.\n\n> Yeah, I tried with the git protocol and it's a few seconds faster (about\n> 14 seconds instead of 17). Ick.\n\nThat's -still- about 13 seconds too much.\n\n> I think it might have something to do with the number of heads we're\n> tracking.\n\nIt really shouldn't matter. You get all the heads in one go with a single \nconnection, so if 32 heads takes 32 times longer, there's something wrong.\n\n> > Also, one thing to try is to just do\n> > \n> > \tstrace -Ttt git-peek-remote ...\n> \n> That's plenty fast, 0.410 seconds, with nothing ugly in the strace.\n\nOk, a \"git fetch\" really shouldn't take any longer than a single \nconnection. However, the fact that you have 32 heads, and it takes pretty \nclose to _exactly_ 32 times 0.410 seconds (32*0.410s = 13.1s) makes me \nsuspect that \"git fetch\" is just broken and fetches one branch at a time. \n\nWhich would be just stupid.\n\nBut look as I might, I see only that one \"git-fetch-pack\" in git-fetch.sh \nthat should trigger. Once. Not 32 times. But your timings sure sound like \nit's doing a _lot_ more than it should.\n\nJunio, any ideas?\n\nKeithp, can you try this trivial patch? It _should_ say something like\n\n\tFetching\n\trefs/heads/master\n\trefs/heads/...\n\trefs/heads/...\n\t...\n\trefs/heads/... from git://..../...\n\nand more importantly, it should say so only once.\n\nAnd then it should leave a \"fetch.trace\" file in your working directory, \nwhich should show where that _one_ thing spends its time.\n\n\t\tLinus\n\n----\ndiff --git a/git-fetch.sh b/git-fetch.sh\nindex 48818f8..4739202 100755\n--- a/git-fetch.sh\n+++ b/git-fetch.sh\n@@ -339,6 +339,8 @@ fetch_main () {\n     ( : subshell because we muck with IFS\n       IFS=\" \t$LF\"\n       (\n+\t  echo \"Fetching $rref from $remote\" >&2\n+\t  strace -o fetch.trace -Ttt \\\n \t  git-fetch-pack $exec $keep --thin \"$remote\" $rref || echo failed \"$remote\"\n       ) |\n       while read sha1 remote_name\n"},{"id":"23172","messageId":"7vsllinj1m.fsf@assigned-by-dhcp.cox.net","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032008590.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-07-04T03:30:29Z","receivedAt":"2006-07-04T03:30:29Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Ok, a \"git fetch\" really shouldn't take any longer than a single \n> connection. However, the fact that you have 32 heads, and it takes pretty \n> close to _exactly_ 32 times 0.410 seconds (32*0.410s = 13.1s) makes me \n> suspect that \"git fetch\" is just broken and fetches one branch at a time. \n>\n> Which would be just stupid.\n>\n> But look as I might, I see only that one \"git-fetch-pack\" in git-fetch.sh \n> that should trigger. Once. Not 32 times. But your timings sure sound like \n> it's doing a _lot_ more than it should.\n>\n> Junio, any ideas?\n\nIsn't that because the repository have 32 subprojects, totally\nunrelated content-wise?  If you have real stuff to pull from\nthere your pack generation needs to do 32 time as much work as\nyou would for a single head in that case.\n\nIf you are discussing \"peek-remote runs, find out the 32 heads\nare all up to date and no pack is generated\" case, then you are\nright.  There is one single fetch-pack to grab the specified\nheads, and after that, an optional single ls-remote and\nfetch-pack runs only once to follow all new tags.\n"},{"id":"23173","messageId":"Pine.LNX.4.64.0607032039010.12404@g5.osdl.org","threadId":"4753","inReplyTo":"7vsllinj1m.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T03:40:43Z","receivedAt":"2006-07-04T03:40:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Junio C Hamano wrote:\n> \n> Isn't that because the repository have 32 subprojects, totally\n> unrelated content-wise?  If you have real stuff to pull from\n> there your pack generation needs to do 32 time as much work as\n> you would for a single head in that case.\n\nNo, Keith said this was for the case where the fetching repository is \nalready totally up-to-date:\n\n    \"And, it's painfully slow, even when the repository is up to date\"\n\nand gave a 17-second time.\n\n\t\t\tLinus\n"},{"id":"23175","messageId":"1151985747.4723.102.camel@neko.keithp.com","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032008590.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-07-04T04:02:27Z","receivedAt":"2006-07-04T04:02:27Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Mon, 2006-07-03 at 20:21 -0700, Linus Torvalds wrote:\n\n> Keithp, can you try this trivial patch? It _should_ say something like\n\nYeah, it says that only once. And, it runs the fetch-pack in about .5\nseconds. And, now the whole process completes in 4.7 seconds; perhaps\nthe remote server is less loaded than earlier this afternoon? It's also\npossible that I was running old git bits here, but I don't think so.\n\n> And then it should leave a \"fetch.trace\" file in your working directory, \n> which should show where that _one_ thing spends its time.\n\nIt looks boring to me and spent 0.55 from start to finish. I can send\nalong the whole trace if you have an acute desire to peer at it.\n\n-- \nkeith.packard@intel.com\n"},{"id":"23176","messageId":"Pine.LNX.4.64.0607032115340.12404@g5.osdl.org","threadId":"4753","inReplyTo":"1151985747.4723.102.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T04:19:49Z","receivedAt":"2006-07-04T04:19:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Keith Packard wrote:\n> \n> Yeah, it says that only once. And, it runs the fetch-pack in about .5\n> seconds. And, now the whole process completes in 4.7 seconds; perhaps\n> the remote server is less loaded than earlier this afternoon?\n\nWell, that's still strange. What takes 4.2 seconds then?\n\n> > And then it should leave a \"fetch.trace\" file in your working directory, \n> > which should show where that _one_ thing spends its time.\n> \n> It looks boring to me and spent 0.55 from start to finish. I can send\n> along the whole trace if you have an acute desire to peer at it.\n\nNo, the 0.5 seconds is what I _expected_. There's something strange going \non in your git fetch that it takes any longer than that.\n\nCan you instrument your \"git-fetch.sh\" script (just add random\n\n\t(echo $LINENO ; date) >&2\n\nlines all over) to see what is so expensive? \n\nThat fetch-pack really should be the most expensive part by far (and half \na second sounds right), but it clearly isn't. At 4.7s, your fetch is still \ntaking about ten times longer than it _should_.\n\n\t\tLinus\n"},{"id":"23177","messageId":"1151987441.4723.110.camel@neko.keithp.com","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032039010.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-07-04T04:30:41Z","receivedAt":"2006-07-04T04:30:41Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Mon, 2006-07-03 at 20:40 -0700, Linus Torvalds wrote:\n\n>     \"And, it's painfully slow, even when the repository is up to date\"\n> \n> and gave a 17-second time.\n\nIt's faster this evening, down to 8 seconds using ssh and 4 seconds\nusing git. I clearly need to force use of the git protocol. Anyone else\nlike the attached patch?\n\n---\n connect.c |   18 ++++++++++++++----\n 1 files changed, 14 insertions(+), 4 deletions(-)\n\ndiff --git a/connect.c b/connect.c\nindex 9a87bd9..e74eddc 100644\n--- a/connect.c\n+++ b/connect.c\n@@ -303,6 +303,7 @@ enum protocol {\n \tPROTO_LOCAL = 1,\n \tPROTO_SSH,\n \tPROTO_GIT,\n+\tPROTO_GIT_SSH,\n };\n \n static enum protocol get_protocol(const char *name)\n@@ -312,9 +313,9 @@ static enum protocol get_protocol(const \n \tif (!strcmp(name, \"git\"))\n \t\treturn PROTO_GIT;\n \tif (!strcmp(name, \"git+ssh\"))\n-\t\treturn PROTO_SSH;\n+\t\treturn PROTO_GIT_SSH;\n \tif (!strcmp(name, \"ssh+git\"))\n-\t\treturn PROTO_SSH;\n+\t\treturn PROTO_GIT_SSH;\n \tdie(\"I don't handle protocol '%s'\", name);\n }\n \n@@ -572,6 +573,14 @@ static void git_proxy_connect(int fd[2],\n \tclose(pipefd[1][0]);\n }\n \n+/* returns whether the specified command can be interpreted by the\ndaemon */\n+int git_is_daemon_command (const char *prog) \n+{\n+\tif (!strcmp(\"git-upload-pack\", prog))\n+\t\treturn 1;\n+\treturn 0;\n+}\n+\n /*\n  * Yeah, yeah, fixme. Need to pass in the heads etc.\n  */\n@@ -641,7 +650,8 @@ int git_connect(int fd[2], char *url, co\n \t\t*ptr = '\\0';\n \t}\n \n-\tif (protocol == PROTO_GIT) {\n+\tif (protocol == PROTO_GIT || \n+\t    (protocol == PROTO_GIT_SSH && git_is_daemon_command (prog))) {\n \t\t/* These underlying connection commands die() if they\n \t\t * cannot connect.\n \t\t */\n@@ -678,7 +688,7 @@ int git_connect(int fd[2], char *url, co\n \t\tclose(pipefd[0][1]);\n \t\tclose(pipefd[1][0]);\n \t\tclose(pipefd[1][1]);\n-\t\tif (protocol == PROTO_SSH) {\n+\t\tif (protocol == PROTO_SSH || protocol == PROTO_GIT_SSH) {\n \t\t\tconst char *ssh, *ssh_basename;\n \t\t\tssh = getenv(\"GIT_SSH\");\n \t\t\tif (!ssh) ssh = \"ssh\";\n-- \n1.4.1.g8fced-dirty\n\n-- \nkeith.packard@intel.com\n"},{"id":"23178","messageId":"1151989503.4723.126.camel@neko.keithp.com","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032115340.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-07-04T05:05:03Z","receivedAt":"2006-07-04T05:05:03Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Mon, 2006-07-03 at 21:19 -0700, Linus Torvalds wrote:\n\n> Can you instrument your \"git-fetch.sh\" script (just add random\n> \n> \t(echo $LINENO ; date) >&2\n> \n> lines all over) to see what is so expensive? \n\n5 Start:                             21:59:01.584648000\n66 After args:                       21:59:01.605987000\n248 fetch_main() start:              21:59:02.408559000\n339 fetch_main() before fetch-pack:  21:59:03.293228000\n387 fetch_main() done:               21:59:04.784388000\n422 After tag following:             21:59:05.311439000\n438 All done:                        21:59:05.315338000\n\nfetch-pack itself took 0.421 seconds (measured with time(1)).\n\nLooks like the bulk of the time here is caused by simple shell\nprocessing overhead, some of which scales with the number of heads and\ntags to track.\n\n-- \nkeith.packard@intel.com\n"},{"id":"23179","messageId":"1151990980.4723.132.camel@neko.keithp.com","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032115340.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-07-04T05:29:40Z","receivedAt":"2006-07-04T05:29:40Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Mon, 2006-07-03 at 21:19 -0700, Linus Torvalds wrote:\n\n> Well, that's still strange. What takes 4.2 seconds then?\n\n$ strace -e trace=execve -f git-fetch 2>&1 | grep execve | sed -e 's/^.*execve(\"//' -e 's/\".*$//' | sort | uniq -c | sort -n\n      1 /bin/rm\n      1 /home/keithp/bin/git\n      1 /home/keithp/bin/git-fetch\n      1 /home/keithp/bin/git-fetch-pack\n      1 /home/keithp/bin/git-ls-remote\n      1 /home/keithp/bin/git-peek-remote\n      1 /usr/bin/sort\n      3 /bin/sed\n      4 /home/keithp/bin/git-repo-config\n     30 /bin/mkdir\n     30 /home/keithp/bin/git-cat-file\n     30 /home/keithp/bin/git-check-ref-format\n     30 /home/keithp/bin/git-merge-base\n     30 /usr/bin/dirname\n     64 /home/keithp/bin/git-rev-parse\n    361 /usr/bin/expr\n\nsomeone sure likes 'expr'...\n\n-- \nkeith.packard@intel.com\n"},{"id":"23180","messageId":"Pine.LNX.4.64.0607032213030.12404@g5.osdl.org","threadId":"4753","inReplyTo":"1151989503.4723.126.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T05:36:15Z","receivedAt":"2006-07-04T05:36:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Keith Packard wrote:\n> \n> 5 Start:                             21:59:01.584648000\n> 66 After args:                       21:59:01.605987000\n> 248 fetch_main() start:              21:59:02.408559000\n> 339 fetch_main() before fetch-pack:  21:59:03.293228000\n> 387 fetch_main() done:               21:59:04.784388000\n> 422 After tag following:             21:59:05.311439000\n> 438 All done:                        21:59:05.315338000\n> \n> fetch-pack itself took 0.421 seconds (measured with time(1)).\n> \n> Looks like the bulk of the time here is caused by simple shell\n> processing overhead, some of which scales with the number of heads and\n> tags to track.\n\nAhh.. Do you have tons of tags at the other end?\n\nLooking closer, I suspect a big part of it is that\n\n\tgit-ls-remote $upload_pack --tags \"$remote\" |\n\tsed -ne 's|^\\([0-9a-f]*\\)[      ]\\(refs/tags/.*\\)^{}$|\\1 \\2|p' |\n\twhile read sha1 name\n\tdo\n\t\t..\n\tdone\n\nloop.\n\nWith a lot of tags, the shell overhead there can indeed be pretty \ndisgusting. And I was wrong - I thought it would do that git-ls-remote \nonly if the first time around we noticed that we would need to, but we do \nactually do it all the time that we're fetching any new branches. \n\nThe sad part is that we really already got the list once, we just never \nsaved it away (ie \"git-fetch-pack\" actually _knows_ what the tags at the \nother end are, and also knows which tags we already have, so if we made \ngit-fetch-pack just create that list and save it off, all the overhead \nwould just go away).\n\nAnd yes, the shell script loops are really really simple, but some of them \nare actually quadratic in the number of refs (O(local*remote)). If this \nwas a C program, we'd never even care, but with shell, the thing is slow \nenough that having even a modest amount of tags and refs is going to just \nmake it waste a lot of time in shell scripting.\n\nWe already do a lot of the infrastructure for \"git fetch\" in C - the \nremotes parsing etc is all things that \"git fetch\" used to share with \"git \npush\", but \"git push\" has been a builtin C program for a while now. I \nsuspect we should just do the same to \"git fetch\", which would make all \nthese issues just totally go away.\n\n\t\t\tLinus\n"},{"id":"23183","messageId":"Pine.LNX.4.64.0607032240260.12404@g5.osdl.org","threadId":"4753","inReplyTo":"1151990980.4723.132.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T05:53:22Z","receivedAt":"2006-07-04T05:53:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Keith Packard wrote:\n>\n>     361 /usr/bin/expr\n> \n> someone sure likes 'expr'...\n\nHeh. That's a very Junio thing to do.\n\nJunio seems to like\n\n\tif expr \"z$string\" : \"z<regexp>\" >/dev/null\n\tthen\n\t\t..\n\nand I think he explained it as being the way old-fashioned users do it.\n\n\t\tLinus\n\t\n"},{"id":"23184","messageId":"7vsllhnb53.fsf@assigned-by-dhcp.cox.net","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032213030.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-07-04T06:21:12Z","receivedAt":"2006-07-04T06:21:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Looking closer, I suspect a big part of it is that\n>\n> \tgit-ls-remote $upload_pack --tags \"$remote\" |\n> \tsed -ne 's|^\\([0-9a-f]*\\)[      ]\\(refs/tags/.*\\)^{}$|\\1 \\2|p' |\n> \twhile read sha1 name\n> \tdo\n> \t\t..\n> \tdone\n>\n> loop.\n\nYes indeed.  Maybe we can do this loop in Perl.  Doing the whole\nthing in C is another option but it would be somewhat painful,\nunless we can deprecate all transport but git native protocols.\n\nOn the other hand, 5 seconds may not matter that much in practice.\n"},{"id":"23189","messageId":"e8d2nu$k44$2@sea.gmane.org","threadId":"4753","inReplyTo":"20060704002138.GB5716@coredump.intra.peff.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-07-04T06:44:48Z","receivedAt":"2006-07-04T06:44:48Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jeff King wrote:\n\n> On Mon, Jul 03, 2006 at 04:14:10PM -0700, Linus Torvalds wrote:\n> \n>> Well, you could use multiple branches in the same repository, even if\nthey \n>> are totally unrealated. That would allow you to fetch them all in one go.\n> \n> One annoying thing about this is that you may want to have several of\n> the branches checked out at a time (i.e., you want the actual directory\n> structure of libXrandr/, Xorg/, etc). You could pull everything down\n> into one repo and point small pseudo-repos at it with alternates, but I\n> would think that would become a mess with pushes. You can do some magic\n> with read-tree --prefix, but again, I'm not sure how you'd make commits\n> on the correct branch.  Is there an easier way to do this?\n\nWrite proper subprojects support for git, or pester someone to write it\n(finally). See Subpro.txt in todo branch.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"23211","messageId":"44AA4CB0.7020604@op5.se","threadId":"4753","inReplyTo":"1151987441.4723.110.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-07-04T11:10:40Z","receivedAt":"2006-07-04T11:10:40Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Keith Packard wrote:\n> On Mon, 2006-07-03 at 20:40 -0700, Linus Torvalds wrote:\n> \n> \n>>    \"And, it's painfully slow, even when the repository is up to date\"\n>>\n>>and gave a 17-second time.\n> \n> \n> It's faster this evening, down to 8 seconds using ssh and 4 seconds\n> using git. I clearly need to force use of the git protocol. Anyone else\n> like the attached patch?\n\nSince it changes the current meaning of ssh+git, I'm not exactly \nthrilled. However, \"git/ssh\" or \"ssh/git\" would work fine for me. The \nslash-separator could be used to say \"fetch over this, push over that\", \nso we can end up with any valid protocol to use for fetches and another \none to push over.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"23214","messageId":"20060704111838.GA4285@spinlock.ch","threadId":"4753","inReplyTo":"44AA4CB0.7020604@op5.se","subject":"Re: git-fetch per-repository speed issues","fromName":"Matthias Kestenholz","fromEmail":"lists@spinlock.ch","sentAt":"2006-07-04T11:18:38Z","receivedAt":"2006-07-04T11:18:38Z","isPatch":false,"sender":{"key":"lists@spinlock.ch","avatar":null},"body":"* Andreas Ericsson (ae@op5.se) wrote:\n> Keith Packard wrote:\n> >On Mon, 2006-07-03 at 20:40 -0700, Linus Torvalds wrote:\n> >\n> >\n> >>   \"And, it's painfully slow, even when the repository is up to date\"\n> >>\n> >>and gave a 17-second time.\n> >\n> >\n> >It's faster this evening, down to 8 seconds using ssh and 4 seconds\n> >using git. I clearly need to force use of the git protocol. Anyone else\n> >like the attached patch?\n> \n> Since it changes the current meaning of ssh+git, I'm not exactly \n> thrilled. However, \"git/ssh\" or \"ssh/git\" would work fine for me. The \n> slash-separator could be used to say \"fetch over this, push over that\", \n> so we can end up with any valid protocol to use for fetches and another \n> one to push over.\n> \n\nIf we would do such a thing, we would be probably better off\nallowing different URLs for pushing and pulling, because the git and\nssh URLs will only be the same, if the git repositories are located\nin the root folder and I suspect that's almost never the case.\n\n\tMatthias\n"},{"id":"23223","messageId":"44AA5987.5060206@op5.se","threadId":"4753","inReplyTo":"20060704111838.GA4285@spinlock.ch","subject":"Re: git-fetch per-repository speed issues","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-07-04T12:05:27Z","receivedAt":"2006-07-04T12:05:27Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Matthias Kestenholz wrote:\n> * Andreas Ericsson (ae@op5.se) wrote:\n> \n>>Keith Packard wrote:\n>>\n>>>On Mon, 2006-07-03 at 20:40 -0700, Linus Torvalds wrote:\n>>>\n>>>\n>>>\n>>>>  \"And, it's painfully slow, even when the repository is up to date\"\n>>>>\n>>>>and gave a 17-second time.\n>>>\n>>>\n>>>It's faster this evening, down to 8 seconds using ssh and 4 seconds\n>>>using git. I clearly need to force use of the git protocol. Anyone else\n>>>like the attached patch?\n>>\n>>Since it changes the current meaning of ssh+git, I'm not exactly \n>>thrilled. However, \"git/ssh\" or \"ssh/git\" would work fine for me. The \n>>slash-separator could be used to say \"fetch over this, push over that\", \n>>so we can end up with any valid protocol to use for fetches and another \n>>one to push over.\n>>\n> \n> \n> If we would do such a thing, we would be probably better off\n> allowing different URLs for pushing and pulling, because the git and\n> ssh URLs will only be the same, if the git repositories are located\n> in the root folder and I suspect that's almost never the case.\n> \n\nTrue. We use relative paths where I work, so for us either way would \nwork. Your way is better though.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"23230","messageId":"e8e28j$v8v$1@sea.gmane.org","threadId":"4753","inReplyTo":"1151949764.4723.51.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-07-04T15:42:12Z","receivedAt":"2006-07-04T15:42:12Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"I wonder if the problem detected here is also responsible with results \nof Jeremy Blosser benchmark comparing git with Mercurial\nhttp://lists.ibiblio.org/pipermail/sm-discuss/2006-May/014586.html\nwhere git wins for clone, status and log, but is slower for pull.\n\nSee summary at\nhttp://git.or.cz/gitwiki/GitBenchmarks#head-85df1bb7f019c4c504e34cde43450ef69349882f\n-- \nJakub Narebski\n"},{"id":"23231","messageId":"20060704163050.GT3305@cip.informatik.uni-erlangen.de","threadId":"4753","inReplyTo":"e8e28j$v8v$1@sea.gmane.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2006-07-04T16:30:50Z","receivedAt":"2006-07-04T16:30:50Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello,\n\n> See summary at\n> http://git.or.cz/gitwiki/GitBenchmarks#head-85df1bb7f019c4c504e34cde43450ef69349882f\n\nthank you for clarifing! I finally understand why Solaris folks prefer\nhg over git: It is dog slow. - So it fits the general philosophy behind\nSolaris.\n\n        Thomas\n"},{"id":"23233","messageId":"7vk66tgt6n.fsf@assigned-by-dhcp.cox.net","threadId":"4753","inReplyTo":"e8e28j$v8v$1@sea.gmane.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-07-04T17:45:36Z","receivedAt":"2006-07-04T17:45:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> I wonder if the problem detected here is also responsible with results \n> of Jeremy Blosser benchmark comparing git with Mercurial\n> http://lists.ibiblio.org/pipermail/sm-discuss/2006-May/014586.html\n> where git wins for clone, status and log, but is slower for pull.\n\nI had an impression, though the report does not talk about this\nspecific detail, that the extra time we are paying is because\nthe \"git pull\" test is done without suppressing the final\ndiffstat phase.\n"},{"id":"23238","messageId":"Pine.LNX.4.64.0607041219540.12404@g5.osdl.org","threadId":"4753","inReplyTo":"7vk66tgt6n.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-04T19:22:05Z","receivedAt":"2006-07-04T19:22:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 4 Jul 2006, Junio C Hamano wrote:\n> \n> I had an impression, though the report does not talk about this\n> specific detail, that the extra time we are paying is because\n> the \"git pull\" test is done without suppressing the final\n> diffstat phase.\n\nI'm pretty sure that was the reason for the particular hg issue. Looking \nat the \"clone\" times, the problem is almost certainly not the actual \npulling.\n\nThe diffstat generation is often the largest part of a git merge. It's \ngotten cheaper since the hg benchmarks were done (I think they were done \nback before the integrated diff generation, so they also have the overhead \nof executing a lot of external GNU diff processes), but it's still not \n\"cheap\".\n\nBut I have to say that the diffstat at least for me is absolutely \ninvaluable.\n\n\t\t\tLinus\n"},{"id":"23244","messageId":"7vfyhhgjxf.fsf@assigned-by-dhcp.cox.net","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607041219540.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-07-04T21:05:32Z","receivedAt":"2006-07-04T21:05:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> But I have to say that the diffstat at least for me is absolutely \n> invaluable.\n\nOh, I absolutely agree with that and somebody who suggests to\nturn it off by default needs a very good argument to convince\nme.\n"},{"id":"23265","messageId":"20060705064749.GA3617@segfault.peff.net","threadId":"4753","inReplyTo":"Pine.LNX.4.64.0607032007290.12404@g5.osdl.org","subject":"Re: git-fetch per-repository speed issues","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-07-05T06:47:49Z","receivedAt":"2006-07-05T06:47:49Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 03, 2006 at 08:07:49PM -0700, Linus Torvalds wrote:\n\n> > Fetching by ssh actually makes two ssh connections (the second is to\n> > grab tags).\n> True. Although that should happen only if there are any new tags.\n\nEither you're wrong or there's a bug in git-fetch. \n\nI think you're missing the call to git-ls-remote --tags to get the list\nof tags (which we will then auto-follow if necessary). So in that case,\nthere would actually be 3 ssh connections. If everything is up to date,\nwe still make 2 connections (one to check refs from remotes file, and\none to check remote tag list).\n\n-Peff\n"},{"id":"23282","messageId":"Pine.LNX.4.64.0607050940210.12404@g5.osdl.org","threadId":"4753","inReplyTo":"20060705064749.GA3617@segfault.peff.net","subject":"Re: git-fetch per-repository speed issues","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-05T16:40:59Z","receivedAt":"2006-07-05T16:40:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Jul 2006, Jeff King wrote:\n> \n> Either you're wrong or there's a bug in git-fetch. \n\nI was wrong - I forgot the git-ls-remote (which really should be \nunnecessary, but the way the git-fetch-pack works, we end up \nre-connecting).\n\n\t\tLinus\n"},{"id":"23338","messageId":"1152228996.22035.6.camel@shinybook.infradead.org","threadId":"4753","inReplyTo":"1151949764.4723.51.camel@neko.keithp.com","subject":"Re: git-fetch per-repository speed issues","fromName":"David Woodhouse","fromEmail":"dwmw2@infradead.org","sentAt":"2006-07-06T23:36:35Z","receivedAt":"2006-07-06T23:36:35Z","isPatch":false,"sender":{"key":"dwmw2@infradead.org","avatar":"https://gravatar.com/avatar/7afd4f07e0cf7d7e046ae2d23678296b37777c96488e6f3451e78a5514154ebd?d=mp&s=160"},"body":"On Mon, 2006-07-03 at 11:02 -0700, Keith Packard wrote:\n>  just uses ssh for everything. This slows down the connection process\n> by several seconds.\n\nOnly if you forgot to use the 'control socket' support, which lets you\nmake a _single_ authenticated connection and re-use it for multiple\nsessions.\n\nhttp://david.woodhou.se/openssh-control.html has a couple of\nimprovements, but the basics are usable in upstream openssh.\n\n-- \ndwmw2\n"}]}