{"thread":{"id":"3298","subject":"Make \"git clone\" less of a deathly quiet experience","startedAt":"2006-02-11T04:31:09Z","lastAt":"2006-02-16T07:33:12Z","messageCount":24,"participants":["Linus Torvalds","Junio C Hamano","Craig Schlenter","Radoslaw Szkodzinski","Petr Baudis","Alex Riesen","Keith Packard","Andreas Ericsson","Martin Langhoff","Eric W. Biederman"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"15900","messageId":"Pine.LNX.4.64.0602102018250.3691@g5.osdl.org","threadId":"3298","inReplyTo":null,"subject":"Make \"git clone\" less of a deathly quiet experience","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-11T04:31:09Z","receivedAt":"2006-02-11T04:31:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nI was on IRC today (which is definitely not normal, but hey, I tried it), \nand somebody was complaining about how horribly slow \"git clone\" was on \nthe WineHQ repository.\n\nThe WineHQ git repo is actually fairly big: 120+MB packed, 220+ thousand \nobjects. So creating the pack is actually a big operation, and yes, it's \ntoo slow. We should be better at it, and it would be good if the pack-file \ngeneration were much faster.\n\nHowever, it turns out that the \"slow\" git-pack-objects was only using up \n2.3% of CPU time. The fact is, the primary reason it took a long time is \nthat even packed, it had to get 120 MB of data. So in this case, it \nappears that the fact that it uses a lot of CPU is actually a good \ntrade-off, because the damn thing would have been even slower if it hadn't \nbeen packed.\n\n(Of course, pre-generated packs would be good regardless)\n\nAnyway, what _really_ made for a pissed-off user was that \"git clone\" was \njust very very silent all the time. No updates on what the hell it was \ndoing. Was it working at all? Was something broken? Is git just a piece of \ncr*p? But \"git clone\" would not say a peep about it.\n\nIt used to be that \"git-unpack-objects\" would give nice percentages, but \nnow that we don't unpack the initial clone pack any more, it doesn't. And \nI'd love to do that nice percentage view in the pack objects downloader \ntoo, but the thing doesn't even read the pack header, much less know how \nmuch it's going to get, so I was lazy and didn't.\n\nInstead, it at least prints out how much data it's gotten, and what the \npackign speed is. Which makes the user realize that it's actually doing \nsomething useful instead of sitting there silently (and if the recipient \nknows how large the final result is, he can at least make a guess about \nwhen it migt be done).\n\nSo with this patch, I get something like this on my DSL line:\n\n\t[torvalds@g5 ~]$ time git clone master.kernel.org:/pub/scm/linux/kernel/git/torvalds/linux-2.6 clone-test\n\tPacking 188543 objects\n\t  48.398MB  (154 kB/s)\n\nwhere even the speed approximation seem sto be roughtly correct (even \nthough my algorithm is a truly stupid one, and only really gives \"speed in \nthe last half second or so\").\n\nAnyway, _something_ like this is definitely needed. It could certainly be \nbetter (if it showed the same kind of thing that git-unpack-objects did, \nthat would be much nicer, but would require parsing the object stream as \nit comes in). But this is  big step forward, I think.\n\nSigned-off-by: Linus Torvalds <torvalds@osdl.org>\n---\n\nComments? Hate-mail? Improvements?\n\ndiff --git a/cache.h b/cache.h\nindex bdbe2d6..c255421 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -348,6 +348,6 @@ extern int copy_fd(int ifd, int ofd);\n \n /* Finish off pack transfer receiving end */\n extern int receive_unpack_pack(int fd[2], const char *me, int quiet);\n-extern int receive_keep_pack(int fd[2], const char *me);\n+extern int receive_keep_pack(int fd[2], const char *me, int quiet);\n \n #endif /* CACHE_H */\ndiff --git a/clone-pack.c b/clone-pack.c\nindex f634431..719e1c4 100644\n--- a/clone-pack.c\n+++ b/clone-pack.c\n@@ -6,6 +6,8 @@ static const char clone_pack_usage[] =\n \"git-clone-pack [--exec=<git-upload-pack>] [<host>:]<directory> [<heads>]*\";\n static const char *exec = \"git-upload-pack\";\n \n+static int quiet = 0;\n+\n static void clone_handshake(int fd[2], struct ref *ref)\n {\n \tunsigned char sha1[20];\n@@ -123,7 +125,9 @@ static int clone_pack(int fd[2], int nr_\n \t}\n \tclone_handshake(fd, refs);\n \n-\tstatus = receive_keep_pack(fd, \"git-clone-pack\");\n+\tif (!quiet)\n+\t\tfprintf(stderr, \"Generating pack ...\\r\");\n+\tstatus = receive_keep_pack(fd, \"git-clone-pack\", quiet);\n \n \tif (!status) {\n \t\tif (nr_match == 0)\n@@ -154,8 +158,10 @@ int main(int argc, char **argv)\n \t\tchar *arg = argv[i];\n \n \t\tif (*arg == '-') {\n-\t\t\tif (!strcmp(\"-q\", arg))\n+\t\t\tif (!strcmp(\"-q\", arg)) {\n+\t\t\t\tquiet = 1;\n \t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tif (!strncmp(\"--exec=\", arg, 7)) {\n \t\t\t\texec = arg + 7;\n \t\t\t\tcontinue;\ndiff --git a/fetch-clone.c b/fetch-clone.c\nindex 859f400..b67d976 100644\n--- a/fetch-clone.c\n+++ b/fetch-clone.c\n@@ -1,6 +1,7 @@\n #include \"cache.h\"\n #include \"exec_cmd.h\"\n #include <sys/wait.h>\n+#include <sys/time.h>\n \n static int finish_pack(const char *pack_tmp_name, const char *me)\n {\n@@ -129,10 +130,12 @@ int receive_unpack_pack(int fd[2], const\n \tdie(\"git-unpack-objects died of unnatural causes %d\", status);\n }\n \n-int receive_keep_pack(int fd[2], const char *me)\n+int receive_keep_pack(int fd[2], const char *me, int quiet)\n {\n \tchar tmpfile[PATH_MAX];\n \tint ofd, ifd;\n+\tunsigned long total;\n+\tstatic struct timeval prev_tv;\n \n \tifd = fd[0];\n \tsnprintf(tmpfile, sizeof(tmpfile),\n@@ -141,6 +144,8 @@ int receive_keep_pack(int fd[2], const c\n \tif (ofd < 0)\n \t\treturn error(\"unable to create temporary file %s\", tmpfile);\n \n+\tgettimeofday(&prev_tv, NULL);\n+\ttotal = 0;\n \twhile (1) {\n \t\tchar buf[8192];\n \t\tssize_t sz, wsz, pos;\n@@ -165,6 +170,27 @@ int receive_keep_pack(int fd[2], const c\n \t\t\t}\n \t\t\tpos += wsz;\n \t\t}\n+\t\ttotal += sz;\n+\t\tif (!quiet) {\n+\t\t\tstatic unsigned long last;\n+\t\t\tstruct timeval tv;\n+\t\t\tunsigned long diff = total - last;\n+\t\t\t/* not really \"msecs\", but a power-of-two millisec (1/1024th of a sec) */\n+\t\t\tunsigned long msecs;\n+\n+\t\t\tgettimeofday(&tv, NULL);\n+\t\t\tmsecs = tv.tv_sec - prev_tv.tv_sec;\n+\t\t\tmsecs <<= 10;\n+\t\t\tmsecs += (int)(tv.tv_usec - prev_tv.tv_usec) >> 10;\n+\t\t\tif (msecs > 500) {\n+\t\t\t\tprev_tv = tv;\n+\t\t\t\tlast = total;\n+\t\t\t\tfprintf(stderr, \"%4lu.%03luMB  (%lu kB/s)        \\r\",\n+\t\t\t\t\ttotal >> 20,\n+\t\t\t\t\t1000*((total >> 10) & 1023)>>10,\n+\t\t\t\t\tdiff / msecs );\n+\t\t\t}\n+\t\t}\n \t}\n \tclose(ofd);\n \treturn finish_pack(tmpfile, me);\ndiff --git a/fetch-pack.c b/fetch-pack.c\nindex 27f5d2a..aa6f42a 100644\n--- a/fetch-pack.c\n+++ b/fetch-pack.c\n@@ -378,7 +378,7 @@ static int fetch_pack(int fd[2], int nr_\n \t\tfprintf(stderr, \"warning: no common commits\\n\");\n \n \tif (keep_pack)\n-\t\tstatus = receive_keep_pack(fd, \"git-fetch-pack\");\n+\t\tstatus = receive_keep_pack(fd, \"git-fetch-pack\", quiet);\n \telse\n \t\tstatus = receive_unpack_pack(fd, \"git-fetch-pack\", quiet);\n \n"},{"id":"15901","messageId":"Pine.LNX.4.64.0602102032410.3691@g5.osdl.org","threadId":"3298","inReplyTo":"Pine.LNX.4.64.0602102018250.3691@g5.osdl.org","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-11T04:37:26Z","receivedAt":"2006-02-11T04:37:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 10 Feb 2006, Linus Torvalds wrote:\n> \n> Instead, it at least prints out how much data it's gotten, and what the \n> packign speed is. Which makes the user realize that it's actually doing \n> something useful instead of sitting there silently (and if the recipient \n> knows how large the final result is, he can at least make a guess about \n> when it migt be done).\n\nBtw, we should print out the other \"stages\" too - the checkout in \nparticular can be a big part of the overhead, and it would probably make \nsense to tell people about the fact that \"hey, now we're checking the \nresult out, we're not actually trying to destroy your disk\".\n\nQuite often, the way to make users happy is not by being impossibly fast \nor beautiful or otherwise wonderful, but by just _managing_ their \nexpectations, so that they don't say \"that's some slow crud\", but instead \nsay \"Ok, it's a nice program, and it's doing a lot of hard work for me\".\n\nIt takes me 15 minutes to clone a kernel repo over the network. Once I can \nsee that most of that is getting a 106MB pack-file at 146 kB/s, I say \"ok, \nthat's fairly reasonable\".\n\n\t\t\tLinus\n"},{"id":"15906","messageId":"7vwtg2o37c.fsf@assigned-by-dhcp.cox.net","threadId":"3298","inReplyTo":"Pine.LNX.4.64.0602102018250.3691@g5.osdl.org","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-11T05:48:55Z","receivedAt":"2006-02-11T05:48:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Anyway, _something_ like this is definitely needed. It could certainly be \n> better (if it showed the same kind of thing that git-unpack-objects did, \n> that would be much nicer, but would require parsing the object stream as \n> it comes in). But this is  big step forward, I think.\n>\n> Signed-off-by: Linus Torvalds <torvalds@osdl.org>\n> ---\n>\n> Comments? Hate-mail? Improvements?\n\nIt probably should default to quiet if (!isatty(1)).\n\nThe real improvement, independent of this client-side patch,\nwould be to reuse recently generated packs, but that needs\nwritable cache directory on the server side.  Another thing that\nI stumbled upon last time I tried it was that it did not look\ntotally trivial to modify the csum-file interface so that I can\nsplice the output from it into two different destinations (one\nto cachefile, the other to the consumer).\n"},{"id":"15907","messageId":"7vslqqo341.fsf@assigned-by-dhcp.cox.net","threadId":"3298","inReplyTo":"Pine.LNX.4.64.0602102032410.3691@g5.osdl.org","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-11T05:50:54Z","receivedAt":"2006-02-11T05:50:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Btw, we should print out the other \"stages\" too - the checkout in \n> particular can be a big part of the overhead, and it would probably make \n> sense to tell people about the fact that \"hey, now we're checking the \n> result out, we're not actually trying to destroy your disk\".\n\nWould you suggest doing that with \"checkout-index -v\", that\nshows \"1 path1\\r2 path2\\r3 path3\\r...\\rDone.\\n\"?\n"},{"id":"15909","messageId":"5C03F8F8-656F-48B0-825C-DE55C837F996@codefountain.com","threadId":"3298","inReplyTo":"7vwtg2o37c.fsf@assigned-by-dhcp.cox.net","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Craig Schlenter","fromEmail":"craig@codefountain.com","sentAt":"2006-02-11T07:35:07Z","receivedAt":"2006-02-11T07:35:07Z","isPatch":false,"sender":{"key":"craig@codefountain.com","avatar":null},"body":"On 11 Feb 2006, at 7:48 AM, Junio C Hamano wrote:\n[snip]\n> The real improvement, independent of this client-side patch,\n> would be to reuse recently generated packs, but that needs\n> writable cache directory on the server side.\n\nSpeaking of improvements, I've noticed that my attempts to track\nthe 2.6 kernel via the git protocol result in inefficiencies from time\nto time when the connection hangs or is terminated when my\nflakey wireless link goes down. When I restart the pull, the data\nthat has already been downloaded is lost and things start from\nscratch which is painful if it's a big update.\n\nIt would be nice if the \"partial pack\" or whatever that has been\ndownloaded at the time of the breakage could be re-used and\nthings could start \"from that point onwards\" or the bits that were\nalready received could be unpacked. Comments?\n\nThank you,\n\n--Craig\n"},{"id":"15910","messageId":"43EDA3D0.7090204@gorzow.mm.pl","threadId":"3298","inReplyTo":"5C03F8F8-656F-48B0-825C-DE55C837F996@codefountain.com","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Radoslaw Szkodzinski","fromEmail":"astralstorm@gorzow.mm.pl","sentAt":"2006-02-11T08:44:00Z","receivedAt":"2006-02-11T08:44:00Z","isPatch":false,"sender":{"key":"astralstorm@gorzow.mm.pl","avatar":null},"body":"Craig Schlenter wrote:\n> On 11 Feb 2006, at 7:48 AM, Junio C Hamano wrote:\n> It would be nice if the \"partial pack\" or whatever that has been\n> downloaded at the time of the breakage could be re-used and\n> things could start \"from that point onwards\" or the bits that were\n> already received could be unpacked. Comments?\n\nIt even already works on plain http repos with git fetch.\n(e.g. WineHQ repository)\nWhy git protocol doesn't support it?\n\n+10\n\n-- \nGPG Key id:  0xD1F10BA2\nFingerprint: 96E2 304A B9C4 949A 10A0  9105 9543 0453 D1F1 0BA2\n\nAstralStorm\n\n"},{"id":"15918","messageId":"20060211130530.GR31278@pasky.or.cz","threadId":"3298","inReplyTo":"43EDA3D0.7090204@gorzow.mm.pl","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-02-11T13:05:30Z","receivedAt":"2006-02-11T13:05:30Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Feb 11, 2006 at 09:44:00AM CET, I got a letter\nwhere Radoslaw Szkodzinski <astralstorm@gorzow.mm.pl> said that...\n> Craig Schlenter wrote:\n> > On 11 Feb 2006, at 7:48 AM, Junio C Hamano wrote:\n> > It would be nice if the \"partial pack\" or whatever that has been\n> > downloaded at the time of the breakage could be re-used and\n> > things could start \"from that point onwards\" or the bits that were\n> > already received could be unpacked. Comments?\n> \n> It even already works on plain http repos with git fetch.\n> (e.g. WineHQ repository)\n> Why git protocol doesn't support it?\n\nBecause it works totally different. When downloading from plain HTTP\nrepos, you are just downloading files from the remote repository and it\nis easy to pick up wherever you left (and last night, I just added a\npossibility to Cogito to resume an interrupted cg-clone by just cd'ing\ninside and running cg-fetch, as is; it's pretty neat) - you just resume\ndownloading of the file you downloaded last, and don't download again\nthe files you already have.\n\nBut the native git protocol works completely differently - you tell the\nserver \"give me all objects you have between object X and head\", the\nobject will generate a completely custom pack just for you and send it\nover the network. The next time you fetch, you just ask for a pack\nbetween object X and head again, but the head can be already totally\ndifferent. What we would have to do is to check for interrupted\npackfiles before fetching, attempt to fix them (cutting out the\nincomplete objects and broken delta chains, if applicable), and then\ntell the remote side to skip those objects; but that may not be easy\nbecause there can be a lot of \"loose fibres\". Another way would be to\njust tell the server \"if head is still Y, start sending the pack only\nafter N bytes\". *shudder*\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nOf the 3 great composers Mozart tells us what it's like to be human,\nBeethoven tells us what it's like to be Beethoven and Bach tells us\nwhat it's like to be the universe.  -- Douglas Adams\n"},{"id":"15920","messageId":"43EDE37C.9050005@gorzow.mm.pl","threadId":"3298","inReplyTo":"20060211130530.GR31278@pasky.or.cz","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Radoslaw Szkodzinski","fromEmail":"astralstorm@gorzow.mm.pl","sentAt":"2006-02-11T13:15:40Z","receivedAt":"2006-02-11T13:15:40Z","isPatch":false,"sender":{"key":"astralstorm@gorzow.mm.pl","avatar":null},"body":"Petr Baudis wrote:\n> But the native git protocol works completely differently - you tell the\n> server \"give me all objects you have between object X and head\", the\n> object will generate a completely custom pack just for you and send it\n> over the network. The next time you fetch, you just ask for a pack\n> between object X and head again, but the head can be already totally\n> different. What we would have to do is to check for interrupted\n> packfiles before fetching, attempt to fix them (cutting out the\n> incomplete objects and broken delta chains, if applicable), and then\n> tell the remote side to skip those objects; but that may not be easy\n> because there can be a lot of \"loose fibres\". Another way would be to\n> just tell the server \"if head is still Y, start sending the pack only\n> after N bytes\". *shudder*\n> \n\nThe other way would be:\n - generate pack file between X and Y\n - start sending from N bytes\n\nIt could break if the repo has been rebased in the meantime.\nBut we could safeguard against it by sending the hash of the packfile\nup to N bytes.\n\n-- \nGPG Key id:  0xD1F10BA2\nFingerprint: 96E2 304A B9C4 949A 10A0  9105 9543 0453 D1F1 0BA2\n\nAstralStorm\n\n"},{"id":"15921","messageId":"20060211133340.GS31278@pasky.or.cz","threadId":"3298","inReplyTo":"7vwtg2o37c.fsf@assigned-by-dhcp.cox.net","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-02-11T13:33:40Z","receivedAt":"2006-02-11T13:33:40Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"BTW, some historical (from the very channel beginning) logs of #git for\nfun, profit and late night reading are available at\nhttp://pasky.or.cz/~pasky/cp/%23git/, e.g. the 2006-02-10 early morning\nfeatures the King Penguin explaining the deepness and intricacies of\npack files construction! Don't miss the opportunity!\n\nNew files won't be world-readable by default, but I hope to get some\nirclogger with cutesy web interface set up for #git.\n\n\nDear diary, on Sat, Feb 11, 2006 at 06:48:55AM CET, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> > Anyway, _something_ like this is definitely needed. It could certainly be \n> > better (if it showed the same kind of thing that git-unpack-objects did, \n> > that would be much nicer, but would require parsing the object stream as \n> > it comes in). But this is  big step forward, I think.\n> >\n> > Signed-off-by: Linus Torvalds <torvalds@osdl.org>\n> > ---\n> >\n> > Comments? Hate-mail? Improvements?\n> \n> It probably should default to quiet if (!isatty(1)).\n\nisatty(2) or something, 1 is in practice always a ref generator. Perhaps\nit would be better not to clutter stderr, though; what about directly\nopening /dev/tty? Does Cygwin support that?\n\n> The real improvement, independent of this client-side patch,\n> would be to reuse recently generated packs, but that needs\n> writable cache directory on the server side.  Another thing that\n> I stumbled upon last time I tried it was that it did not look\n> totally trivial to modify the csum-file interface so that I can\n> splice the output from it into two different destinations (one\n> to cachefile, the other to the consumer).\n\nYes, I said that on IRC yesterday as well. I don't think even a cache is\nneeded; just look at the repository and say:\n\n\t* while there are packs containing only objects we are going to\n\t  send, pick the largest one and send it as-is.\n\t* if there is a pack with more than a 75% (totally arbitrary)\n\t  overlap with the objects we are going to send, send it as-is.\n\t* pack the loose objects.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nOf the 3 great composers Mozart tells us what it's like to be human,\nBeethoven tells us what it's like to be Beethoven and Bach tells us\nwhat it's like to be the universe.  -- Douglas Adams\n"},{"id":"15922","messageId":"20060211134142.GT31278@pasky.or.cz","threadId":"3298","inReplyTo":"20060211133340.GS31278@pasky.or.cz","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-02-11T13:41:42Z","receivedAt":"2006-02-11T13:41:42Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sat, Feb 11, 2006 at 02:33:40PM CET, I got a letter\nwhere Petr Baudis <pasky@suse.cz> said that...\n> BTW, some historical (from the very channel beginning) logs of #git for\n> fun, profit and late night reading are available at\n> http://pasky.or.cz/~pasky/cp/%23git/, e.g. the 2006-02-10 early morning\n> features the King Penguin explaining the deepness and intricacies of\n> pack files construction! Don't miss the opportunity!\n> \n> New files won't be world-readable by default, but I hope to get some\n> irclogger with cutesy web interface set up for #git.\n\nLike,\n\n\thttp://colabti.de/irclogger/irclogger_logs/git\n\n(Courtesy of Francois Beerten.)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nOf the 3 great composers Mozart tells us what it's like to be human,\nBeethoven tells us what it's like to be Beethoven and Bach tells us\nwhat it's like to be the universe.  -- Douglas Adams\n"},{"id":"15928","messageId":"20060211172403.GA10099@steel.home","threadId":"3298","inReplyTo":"20060211133340.GS31278@pasky.or.cz","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-02-11T17:24:03Z","receivedAt":"2006-02-11T17:24:03Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Petr Baudis, Sat, Feb 11, 2006 14:33:40 +0100:\n> > It probably should default to quiet if (!isatty(1)).\n> \n> isatty(2) or something, 1 is in practice always a ref generator. Perhaps\n> it would be better not to clutter stderr, though; what about directly\n> opening /dev/tty? Does Cygwin support that?\n\nIt can't. Windows has no terminals (as in \"none at all\"). It has a\nConsole, which is a special kind of window attached to an application\nand where the unbuffered stdout and stderr are magically redirected.\n\nA test for is stdout/err is a tty can only check if the process has\nthe console attached, and an attempt to open it for writing will\nprobably just create the thing.\n"},{"id":"15929","messageId":"Pine.LNX.4.64.0602110936510.3691@g5.osdl.org","threadId":"3298","inReplyTo":"7vslqqo341.fsf@assigned-by-dhcp.cox.net","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-11T17:39:38Z","receivedAt":"2006-02-11T17:39:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 10 Feb 2006, Junio C Hamano wrote:\n> \n> Would you suggest doing that with \"checkout-index -v\", that\n> shows \"1 path1\\r2 path2\\r3 path3\\r...\\rDone.\\n\"?\n\nNot if it shows every single path.\n\nWhen going tty output, we should be careful to limit it to not do tons and \ntons of lines. The download output does gettimeofday to limit itself to \nmax 2 times per sec, and the percentage output of git-unpack-objects \nsimilarly limits itself so that it never spews _tons_ of stuff to the \nterminal.\n\nUnder many loads, the terminal will be a lot slower than actually writing \na file (\"context switch to gnome-term + context switch to X + set up \ncomplex text output\").\n\n\t\tLinus\n"},{"id":"15930","messageId":"Pine.LNX.4.64.0602110943170.3691@g5.osdl.org","threadId":"3298","inReplyTo":"7vwtg2o37c.fsf@assigned-by-dhcp.cox.net","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-11T17:45:56Z","receivedAt":"2006-02-11T17:45:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 10 Feb 2006, Junio C Hamano wrote:\n> \n> It probably should default to quiet if (!isatty(1)).\n\nSounds fine. isatty(2), though, since we use stderr for these messages \n(stdout is usually the data-stream).\n\n> The real improvement, independent of this client-side patch,\n> would be to reuse recently generated packs, but that needs\n> writable cache directory on the server side.\n\nMore importantly, it really wouldn't have helped that much in this \nsituation. At least for me, the network is 90% of the problem, the \npack-file generation is at most 10%. So cached packfiles really only \nmatter for server-side problems (high CPU load, or lack of memory, or \nheavy disk activity).\n\nSo the problems really are very independent.\n\n\t\t\tLinus\n"},{"id":"15932","messageId":"20060211183959.GA9984@steel.home","threadId":"3298","inReplyTo":"Pine.LNX.4.64.0602102018250.3691@g5.osdl.org","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-02-11T18:39:59Z","receivedAt":"2006-02-11T18:39:59Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"Linus Torvalds, Sat, Feb 11, 2006 05:31:09 +0100:\n> So with this patch, I get something like this on my DSL line:\n> \n> \t[torvalds@g5 ~]$ time git clone master.kernel.org:/pub/scm/linux/kernel/git/torvalds/linux-2.6 clone-test\n> \tPacking 188543 objects\n> \t  48.398MB  (154 kB/s)\n\nI get this:\n\n    $ git clone . ../cloned\n    Packing 15440 objects\n    $ 2 kB/s)\n\nI'd put a \\n before finish_pack to make it nicer.\n\nSigned-off-by: Alex Riesen <raa.lkml@gmail.com>\n\ndiff --git a/fetch-clone.c b/fetch-clone.c\nindex b67d976..37141e9 100644\n--- a/fetch-clone.c\n+++ b/fetch-clone.c\n@@ -193,5 +193,7 @@ int receive_keep_pack(int fd[2], const c\n \t\t}\n \t}\n \tclose(ofd);\n+\tif ( !quiet )\n+\t    fputc('\\n', stderr);\n \treturn finish_pack(tmpfile, me);\n }\n"},{"id":"15935","messageId":"Pine.LNX.4.64.0602111100270.3691@g5.osdl.org","threadId":"3298","inReplyTo":"20060211183959.GA9984@steel.home","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-02-11T19:04:35Z","receivedAt":"2006-02-11T19:04:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 11 Feb 2006, Alex Riesen wrote:\n> \n> I'd put a \\n before finish_pack to make it nicer.\n\nYes.\n\nDuh. I did all my testing with \"time git clone ..\", so I had the extra \\n \nadded by the fact that \"time\" itself will do it.\n\nSide comment: the pack preparation stage seems to take about 90s for the \nkernel. Of course, that will keep growing with history, but so will \nprobably the pack-size, so percentage-wise, the 90% / 10% thing is likely \nto hold for DSL (yes, DSL gets faster too, but so do CPU ;).\n\nThat 90s is unquestionably irritating, though, so we do want to either \ncache them, or add similar \"I'm working on it\" output to that phase too.\n\n\t\tLinus\n"},{"id":"15936","messageId":"1139685031.4183.31.camel@evo.keithp.com","threadId":"3298","inReplyTo":"Pine.LNX.4.64.0602110943170.3691@g5.osdl.org","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-02-11T19:10:29Z","receivedAt":"2006-02-11T19:10:29Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Sat, 2006-02-11 at 09:45 -0800, Linus Torvalds wrote:\n\n> More importantly, it really wouldn't have helped that much in this \n> situation. At least for me, the network is 90% of the problem, the \n> pack-file generation is at most 10%. So cached packfiles really only \n> matter for server-side problems (high CPU load, or lack of memory, or \n> heavy disk activity).\n\nI'd like to see git use less CPU than CVS does on my distribution host;\nsome mechanism for re-using either existing or cached packs would help a\nwhole lot with that. The alternative is to see people switch to rsync\ninstead, which seems like a far worse idea.   \n\n-- \nkeith.packard@intel.com\n"},{"id":"15954","messageId":"43EEAEF3.7040202@op5.se","threadId":"3298","inReplyTo":"1139685031.4183.31.camel@evo.keithp.com","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-02-12T03:43:47Z","receivedAt":"2006-02-12T03:43:47Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Keith Packard wrote:\n> On Sat, 2006-02-11 at 09:45 -0800, Linus Torvalds wrote:\n> \n> \n>>More importantly, it really wouldn't have helped that much in this \n>>situation. At least for me, the network is 90% of the problem, the \n>>pack-file generation is at most 10%. So cached packfiles really only \n>>matter for server-side problems (high CPU load, or lack of memory, or \n>>heavy disk activity).\n> \n> \n> I'd like to see git use less CPU than CVS does on my distribution host;\n> some mechanism for re-using either existing or cached packs would help a\n> whole lot with that. The alternative is to see people switch to rsync\n> instead, which seems like a far worse idea.   \n> \n\nA weird oddity; Cloning is faster over rsync, day-to-day pulling is not.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"15957","messageId":"1139717510.4183.34.camel@evo.keithp.com","threadId":"3298","inReplyTo":"43EEAEF3.7040202@op5.se","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-02-12T04:11:49Z","receivedAt":"2006-02-12T04:11:49Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Sun, 2006-02-12 at 04:43 +0100, Andreas Ericsson wrote:\n\n> A weird oddity; Cloning is faster over rsync, day-to-day pulling is not.\n\nPrecisely. If the protocol could deliver existing packs instead of\nunpacking and repacking them, then git would be as fast as rsync and I\nwouldn't have to worry about supporting two protocols.\n\n-- \nkeith.packard@intel.com\n"},{"id":"15966","messageId":"43EF15D1.1050609@op5.se","threadId":"3298","inReplyTo":"1139717510.4183.34.camel@evo.keithp.com","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-02-12T11:02:41Z","receivedAt":"2006-02-12T11:02:41Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Keith Packard wrote:\n> On Sun, 2006-02-12 at 04:43 +0100, Andreas Ericsson wrote:\n> \n> \n>>A weird oddity; Cloning is faster over rsync, day-to-day pulling is not.\n> \n> \n> Precisely. If the protocol could deliver existing packs instead of\n> unpacking and repacking them, then git would be as fast as rsync and I\n> wouldn't have to worry about supporting two protocols.\n> \n\nCaching features have been discussed, but that means the daemon needs to \nhave write-access to some directory within the repository. It would also \nwork poorly for projects that see very rapid development unless the \ncached pack-files can be amended to. A sort of \"create packs on demand\". \nIt shouldn't be too difficult, really.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"16064","messageId":"1139778267.4183.66.camel@evo.keithp.com","threadId":"3298","inReplyTo":"43EF15D1.1050609@op5.se","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-02-12T21:04:27Z","receivedAt":"2006-02-12T21:04:27Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Sun, 2006-02-12 at 12:02 +0100, Andreas Ericsson wrote:\n\n> Caching features have been discussed, but that means the daemon needs to \n> have write-access to some directory within the repository. \n\nCaching seems a bit dicey to me; security concerns and all. I would much\nrather have it discover packs on disk that provided a subset of the\nnecessary objects; repository cloning would then be a process of\ndelivering any available packs and then packing up the remaining\nobjects. Clever administration of the repository could then construct a\nsingle pack of 'historical' data followed by periodic packs of\nincremental data.\n\nYeah, I know, I should just implement this and see how well it works in\npractice. I apologize for thinking in public.     \n      \n-- \nkeith.packard@intel.com\n"},{"id":"16011","messageId":"46a038f90602121806jfcaac41tb98b8b4cd4c07c23@mail.gmail.com","threadId":"3298","inReplyTo":"1139717510.4183.34.camel@evo.keithp.com","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-02-13T02:06:42Z","receivedAt":"2006-02-13T02:06:42Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 2/12/06, Keith Packard <keithp@keithp.com> wrote:\n> On Sun, 2006-02-12 at 04:43 +0100, Andreas Ericsson wrote:\n>\n> > A weird oddity; Cloning is faster over rsync, day-to-day pulling is not.\n>\n> Precisely. If the protocol could deliver existing packs instead of\n> unpacking and repacking them, then git would be as fast as rsync and I\n> wouldn't have to worry about supporting two protocols.\n\n+1... there should be an easy-to-compute threshold trigger to say --\nhey, let's quit being smart and send this client the packs we got and\nget it over with. Or perhaps a client flag so large projects can\nrecommend that uses do their initial clone with --gimme-all-packs?\n\nMy workaround for large repos is to clone over http, and s/http:/git:/\non the origin file once it's done ;-)\n\n\nmartin\n"},{"id":"16015","messageId":"7v4q3453qu.fsf@assigned-by-dhcp.cox.net","threadId":"3298","inReplyTo":"46a038f90602121806jfcaac41tb98b8b4cd4c07c23@mail.gmail.com","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-13T03:36:41Z","receivedAt":"2006-02-13T03:36:41Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> writes:\n\n> +1... there should be an easy-to-compute threshold trigger to say --\n> hey, let's quit being smart and send this client the packs we got and\n> get it over with. Or perhaps a client flag so large projects can\n> recommend that uses do their initial clone with --gimme-all-packs?\n\nWhat upload-pack does boils down to:\n\n    * find out the latest of what client has and what client asked.\n\n    * run \"rev-list --objects ^client ours\" to make a list of\n      objects client needs.  The actual command line has multiple\n      \"clients\" to exclude what is unneeded to be sent, and\n      multiple \"ours\" to include refs asked.  When you are doing\n      a full clone, ^client is empty and ours is essentially\n      --all.\n\n    * feed that output to \"pack-objects --stdout\" and send out\n      the result.\n\nIf you run this command:\n\n\t$ git-rev-list --objects --all |\n          git-pack-objects --stdout >/dev/null \n\nIt would say some things.  The phases of operations are:\n\n\tGenerating pack...\n\tCounting objects XXXX...\n        Done counting XXXX objects.\n        Packing XXXXX objects.....\n\nPhase (1).  Between the time it says \"Generating pack...\" upto\n\"Done counting XXXX objects.\", the time is spent by rev-list to\nlist up all the objects to be sent out.\n\nPhase (2). After that, it tries to make decision what object to\ndelta against what other object, while twenty or so dots are\nprinted after \"Packing XXXXX objects.\" (see #git irc log a\ncouple of days ago; Linus describes how pack building works).\n\nPhase (3). After the dot stops, the program becomes silent.\nThat is where it actually does delta compression and writeout.\n\nYou would notice that quite a lot of time is spent in all\nphases.\n\nThere is an internal hook to create full repository pack inside\nupload-pack (which is what runs on the other end when you run\nfetch-pack or clone-pack), but it works slightly differently\nfrom what you are suggesting, in that it still tries to do the\n\"correct\" thing.  It still runs \"rev-list --objects --all\", so\n\"dangling objects\" are never sent out.\n\nWe could cheat in all phases to speed things up, at the expense\nof ending up sending excess objects.  So let's pretend we\ndecided to treat everything in .git/objects/packs/pack-* (and\nthe ones found in alternates as well) have interesting objects\nfor the cloner.\n\n(1) This part unfortunately cannot be totally eliminated.  By\n    assume all packs are interesting, we could use the object\n    names from the pack index, which is a lot cheaper than\n    rev-list object traversal.  We still need to run rev-list\n    --objects --all --unpacked to pick up loose objects we would\n    not be able to tell by looking at the pack index to cover\n    the rest.\n\n    This however needs to be done in conjunction with the second\n    phase change.  pack-objects depends on the hint rev-list\n    --objects output gives it to group the blobs and trees with\n    the same pathnames together, and that greatly affects the\n    packing efficiency.  Unfortunately pack index does not have\n    that information -- it does not know type, nor pathnames.\n    Type is relatively cheap to obtain but pathnames for blob\n    objects are inherently unavailable.\n\n(2) This part can be mostly eliminated for already packed\n    objects, because we have already decided to cheat by sending\n    everything, so we can just reuse how objects are deltified\n    in existing packs.  It still needs to be done for loose\n    objects we collected to fill the gap in (1).\n\n(3) This also can be sped up by reusing what are already in\n    packs.  Pack index records starting (but not end) offset of\n    each object in the pack, so we can sort by offset to find\n    out which part of the existing pack corresponds to what\n    object, to reorder the objects in the final pack.  This\n    needs to be done somewhat carefully to preserve the locality\n    of objects (again, see #git log).  The deltifying and\n    compressing for loose objects cannot be avoided.\n\n    While we are writing things out in (3), we need to keep\n    track of running SHA1 sum of what we write out so that we\n    can fill out the correct checksum at the end, but I am\n    guessing that is relatively cheap compared to the\n    deltification and compression cost we are currently paying\n    in this phase.\n\nNB. In the #git log, Linus made it sound like I am clueless\nabout how pack is generated, but if you check commit 9d5ab96,\nthe \"recency of delta is inherited from base\", one of the tricks\nthat have a big performance impact, was done by me ;-).\n"},{"id":"16249","messageId":"m1ek23rduh.fsf@ebiederm.dsl.xmission.com","threadId":"3298","inReplyTo":"43EF15D1.1050609@op5.se","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2006-02-16T06:56:38Z","receivedAt":"2006-02-16T06:56:38Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Andreas Ericsson <ae@op5.se> writes:\n\n> Keith Packard wrote:\n>> On Sun, 2006-02-12 at 04:43 +0100, Andreas Ericsson wrote:\n>>\n>>>A weird oddity; Cloning is faster over rsync, day-to-day pulling is not.\n>> Precisely. If the protocol could deliver existing packs instead of\n>> unpacking and repacking them, then git would be as fast as rsync and I\n>> wouldn't have to worry about supporting two protocols.\n>>\n>\n> Caching features have been discussed, but that means the daemon needs to have\n> write-access to some directory within the repository. It would also work poorly\n> for projects that see very rapid development unless the cached pack-files can be\n> amended to. A sort of \"create packs on demand\". It shouldn't be too difficult,\n> really.\n\nActually for the clone case we don't need a writable directory for the\ngit-daemon. \n\nIf we assume that a repository up for download is reasonably packed,\nwe can just lob all of the packs in the current repository, and then\npack the few remaining objects and send them.\n\nI don't know how well multiple packs will work with the current git\nprotocol but it should be pretty natural, and the clone case is easy\ndetect as there are no heads in common.  Can that be detected quickly?\n\nI don't have a patch but it feels like a pretty straight forward thing\nto implement.\n\nEric\n"},{"id":"16250","messageId":"7v7j7vhi6f.fsf@assigned-by-dhcp.cox.net","threadId":"3298","inReplyTo":"m1ek23rduh.fsf@ebiederm.dsl.xmission.com","subject":"Re: Make \"git clone\" less of a deathly quiet experience","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-02-16T07:33:12Z","receivedAt":"2006-02-16T07:33:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"ebiederm@xmission.com (Eric W. Biederman) writes:\n\n> I don't know how well multiple packs will work with the current git\n> protocol...\n\nThen I wonder why you are making this observation ... ;-)\n\nIn any case, I suspect this would be helped to a certain degree\nby the pack-object that reuses delta data from existing packs,\nif your repository is reasonably packed.\n"}]}