{"thread":{"id":"4219","subject":"irc usage..","startedAt":"2006-05-20T17:26:22Z","lastAt":"2006-06-05T16:07:43Z","messageCount":82,"participants":["Linus Torvalds","Junio C Hamano","Jakub Narebski","Yann Dirson","Donnie Berkholz","Thomas Glanzmann","Martin Langhoff","Matthias Lederhofer","Matthias Urlichs","Jeff King","Morten Welinder","Alec Warner","Sean"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"20329","messageId":"Pine.LNX.4.64.0605201016090.10823@g5.osdl.org","threadId":"4219","inReplyTo":null,"subject":"irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-20T17:26:22Z","receivedAt":"2006-05-20T17:26:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nI hate irc.\n\nI'm reading the irc logs, and seeing that people have problems, but (a) it \nwas while I was asleep and (b) irc use doesn't encourage people to \nactually explain what the problems _are_, so I have no clue.\n\nSo now I know that \"spyderous\" has problems importing some 1GB gentoo CVS \narchive, but that's pretty much it. Grr.\n\nAre people afraid to post to git@vger.kernel.org, or what?\n\nI saw that people tried to suggest posting to the git mailing list, but \ncan any of you who are active on irc be a bit more forceful? And perhaps \nwe don't make this mailing list address well enough known? \n\nAs far as I'm aware, the git mailing list isn't closed, so people should \nbe able to post here without even subscribing. I can well understand that \nyou might not want to subscribe and prefer to look ove rthe list through \nsome archive setup (the way I look at the irc logs), and maybe we should \njust make the git mailing list address more obvious.\n\nRight now, the \"community\" page at http://git.or.cz/community.html doesn't \neven mention the git mailing list address directly, it just tells you how \nyou can subscribe and read the archives.\n\nCan we perhaps fix that, and the people who are active on irc please also \nmake it clear to people that if they have some real problems that don't \nget an immediate answer, the git mailing list ends up where a lot of \npeople can actually look more closely at it.. And tell them what the \naddress is.\n\n\t\t\tLinus\n"},{"id":"20331","messageId":"7vzmhc1t70.fsf@assigned-by-dhcp.cox.net","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605201016090.10823@g5.osdl.org","subject":"Re: irc usage..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-20T17:50:43Z","receivedAt":"2006-05-20T17:50:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> I hate irc.\n>...\n> Can we perhaps fix that, and the people who are active on irc please also \n> make it clear to people that if they have some real problems that don't \n> get an immediate answer, the git mailing list ends up where a lot of \n> people can actually look more closely at it.. And tell them what the \n> address is.\n\nI hate irc, too.  Number of times easily solvable usage problems\ncome up and I look at the log to realize when the solutions\nsuggested were waaaaay suboptimal it is too late (with loops\nbeing quite active recently things have improved a lot, but we\nshould not expect him to be 24/7).\n\nMaybe somebody can run a dumb 'bot that notices somebody said\nsomething that ends with a '?' and there is no activity there\nfor N minutes and inject a recorded message that reminds the\nmailing list address ;-).\n"},{"id":"20335","messageId":"e4noha$p92$1@sea.gmane.org","threadId":"4219","inReplyTo":"7vzmhc1t70.fsf@assigned-by-dhcp.cox.net","subject":"Re: irc usage..","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-20T18:52:46Z","receivedAt":"2006-05-20T18:52:46Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n\n> Maybe somebody can run a dumb 'bot that notices somebody said\n> something that ends with a '?' and there is no activity there\n> for N minutes and inject a recorded message that reminds the\n> mailing list address ;-).\n\nOr something like fsbot or other bots on #emacs channel\n\n   http://www.emacswiki.org/cgi-bin/wiki/EmacsChannel\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"20337","messageId":"20060520203911.GI6535@nowhere.earth","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605201016090.10823@g5.osdl.org","subject":"Re: irc usage..","fromName":"Yann Dirson","fromEmail":"ydirson@altern.org","sentAt":"2006-05-20T20:39:11Z","receivedAt":"2006-05-20T20:39:11Z","isPatch":false,"sender":{"key":"ydirson@altern.org","avatar":"https://avatars.githubusercontent.com/u/1190950?v=4"},"body":"On Sat, May 20, 2006 at 10:26:22AM -0700, Linus Torvalds wrote:\n> I'm reading the irc logs, and seeing that people have problems, but (a) it \n> was while I was asleep and (b) irc use doesn't encourage people to \n> actually explain what the problems _are_, so I have no clue.\n> \n> So now I know that \"spyderous\" has problems importing some 1GB gentoo CVS \n> archive, but that's pretty much it. Grr.\n\nFWIW, I have mentionned a problem that may be the same, under\nMessage-ID <20060107090148.GB32585@nowhere.earth>, that was on January\n7th.  Namely, when importing a repository with very large files over\npserver or ssh, timeouts can occur and prevent the import from\nworking.  But, as you said, it's not easy to get precise info from the\nlogs :)\n\nBest regards,\n-- \nYann Dirson    <ydirson@altern.org> |\nDebian-related: <dirson@debian.org> |   Support Debian GNU/Linux:\n                                    |  Freedom, Power, Stability, Gratis\n     http://ydirson.free.fr/        | Check <http://www.debian.org/>\n"},{"id":"20342","messageId":"446F95A2.6040909@gentoo.org","threadId":"4219","inReplyTo":"20060520203911.GI6535@nowhere.earth","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-20T22:18:10Z","receivedAt":"2006-05-20T22:18:10Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Yann Dirson wrote:\n> On Sat, May 20, 2006 at 10:26:22AM -0700, Linus Torvalds wrote:\n>> I'm reading the irc logs, and seeing that people have problems, but (a) it \n>> was while I was asleep and (b) irc use doesn't encourage people to \n>> actually explain what the problems _are_, so I have no clue.\n>>\n>> So now I know that \"spyderous\" has problems importing some 1GB gentoo CVS \n>> archive, but that's pretty much it. Grr.\n\nHi all,\n\nI just subscribed and this post is the only one I've got from the\nthread, so I'm responding to it instead of the original. Gentoo's an\nIRC-based community, so I tend to try IRC first for any problems I have\nand fall back to the list later if I can't get things figured out.\n\nHere's a rough summary:\n\nOur main repo is actually a bit over 2G (2103621223) now that I check,\nbut it's not very complex. There's actually just one branch, and I don't\nthink anyone would care if we lost the history from it because it's a\nrelease branch from a few years ago.\n\nSomebody else tried importing it with git-cvsimport, but he said he hit\nsome kind of problem and recalled that it was a cvsps segfault. Sounds\nabout right, since I've never gotten cvsps to run successfully on the\nwhole repo either.\n\nI tried with parsecvs, but it runs into OOM even on a machine with 4G\nRAM after reading in all the ,v files, presumably while it's building\nsome huge tree of changesets in memory. Keith Packard's suggested that\nthere are ways to reduce parsecvs's memory use, because it retains the\nfull tree in memory for each revision rather than just the files that\nactually changed. But my C skills are pretty weak; I'm an OK reader but\nnot much of a writer yet.\n\nThanks,\nDonnie\n\n"},{"id":"20343","messageId":"Pine.LNX.4.64.0605201543260.3649@g5.osdl.org","threadId":"4219","inReplyTo":"446F95A2.6040909@gentoo.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-20T22:45:00Z","receivedAt":"2006-05-20T22:45:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 20 May 2006, Donnie Berkholz wrote:\n> \n> Our main repo is actually a bit over 2G (2103621223) now that I check,\n> but it's not very complex. There's actually just one branch, and I don't\n> think anyone would care if we lost the history from it because it's a\n> release branch from a few years ago.\n\nCan you point to it? I'm not a CVS user, but I've played with cvsps before \n(to get it to work), and I'm a humanitarian - rescuing people from CVS is \nto me not just a good idea, it's a moral imperative.\n\n\t\tLinus\n"},{"id":"20349","messageId":"446FA262.7080900@gentoo.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605201543260.3649@g5.osdl.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-20T23:12:34Z","receivedAt":"2006-05-20T23:12:34Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Sat, 20 May 2006, Donnie Berkholz wrote:\n>> Our main repo is actually a bit over 2G (2103621223) now that I check,\n>> but it's not very complex. There's actually just one branch, and I don't\n>> think anyone would care if we lost the history from it because it's a\n>> release branch from a few years ago.\n> \n> Can you point to it? I'm not a CVS user, but I've played with cvsps before \n> (to get it to work), and I'm a humanitarian - rescuing people from CVS is \n> to me not just a good idea, it's a moral imperative.\n\nI don't want to post the link publicly for a few reasons, including the\nhuge amount of bandwidth it would suck up for lots of people to download\nit. I've sent it to you off-list, and if anyone else would also like it,\nplease drop me a note.\n\nThanks,\nDonnie\n\n"},{"id":"20362","messageId":"446FBEE3.2060809@gentoo.org","threadId":"4219","inReplyTo":"446F95A2.6040909@gentoo.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-21T01:14:11Z","receivedAt":"2006-05-21T01:14:11Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Donnie Berkholz wrote:\n> Somebody else tried importing it with git-cvsimport, but he said he hit\n> some kind of problem and recalled that it was a cvsps segfault. Sounds\n> about right, since I've never gotten cvsps to run successfully on the\n> whole repo either.\n\nMuch to my surprise, a cvsps run I started earlier has just finished\nwithout segfaulting. But attempts to actually run cvsps (e.g., cvsps -a\nspyderous) spit thousands of warnings of \"WARNING: revision 1.1.1.1 of\nfile $FILENAME on unnamed branch\".\n\nThanks,\nDonnie\n\n"},{"id":"20391","messageId":"20060521094606.GD5545@cip.informatik.uni-erlangen.de","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605201543260.3649@g5.osdl.org","subject":"Re: irc usage..","fromName":"Thomas Glanzmann","fromEmail":"sithglan@stud.uni-erlangen.de","sentAt":"2006-05-21T09:46:06Z","receivedAt":"2006-05-21T09:46:06Z","isPatch":false,"sender":{"key":"sithglan@stud.uni-erlangen.de","avatar":null},"body":"Hello Linus,\n\n> and I'm a humanitarian - rescuing people from CVS is \n> to me not just a good idea, it's a moral imperative.\n\nyou're a very brave man.\n\n        Thomas\n"},{"id":"20402","messageId":"Pine.LNX.4.64.0605211209080.3649@g5.osdl.org","threadId":"4219","inReplyTo":"446FA262.7080900@gentoo.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-21T19:24:16Z","receivedAt":"2006-05-21T19:24:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 20 May 2006, Donnie Berkholz wrote:\n> \n> I don't want to post the link publicly for a few reasons, including the\n> huge amount of bandwidth it would suck up for lots of people to download\n> it. I've sent it to you off-list, and if anyone else would also like it,\n> please drop me a note.\n\nOk. It's still converting (that's a big archive), but it has passed the \ncvsps stage without errors for me, and the conversion so far seems ok. But \nit has only gotten to \n\n\tAuthor: vapier <vapier>  2002-09-23 12:32:42\n\tChanged GPL to GPL-2 in LICENSE and updated SRC_URI to use mirror:\n\nso it has converted only slightly more than the first two years of \nhistory in the roughly 30 minutes I've let it run. So it will take several \nhours.\n\nThe reason it works for me is likely simply the fact that I had a few \npatches to my cvsps already. I'm appending the stupid patches, I'm not \nguaranteeing that they are correct at all, although the three _committed_ \npatches are almost certainly correct (and the last uncommitted one is \nalmost certainly totally broken). The patches are against clean cvsps 2.1.\n\nAlso, when I say \"the conversion so far seems ok\", I obviously don't \nactually know what the hell the archive is supposed to look like, so I can \nonly say that the end result seems not totally insane.\n\nTo do a good conversion, you'll want to make sure that you have a author \nname conversion file. See the \"-A\" flag in \"git help cvsimport\" (if you \nhave the man-pages installed).\n\n\t\tLinus\n\n---\ncommit 534120d9a47062eecd7b53fd7ac0b70d97feb4fd\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Wed Mar 22 11:20:59 2006 -0800\n\n    Increase log-length limit to 64kB\n    \n    Yeah, it should be dynamic. I'm lazy.\n---\n cvsps_types.h |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/cvsps_types.h b/cvsps_types.h\nindex b41e2a9..dba145d 100644\n--- a/cvsps_types.h\n+++ b/cvsps_types.h\n@@ -8,7 +8,7 @@ #define CVSPS_TYPES_H\n \n #include <time.h>\n \n-#define LOG_STR_MAX 32768\n+#define LOG_STR_MAX 65536\n #define AUTH_STR_MAX 64\n #define REV_STR_MAX 64\n #define MIN(a, b) ((a) < (b) ? (a) : (b))\n\n\ncommit 82fcf7e31bbeae3b01a8656549e9b8fd89d598eb\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Wed Mar 22 11:23:37 2006 -0800\n\n    Improve handling of file collisions in the same patchset\n    \n    Take the file revision into account.\n---\n cvsps.c |   27 +++++++++++++++++++++++++--\n 1 files changed, 25 insertions(+), 2 deletions(-)\n\ndiff --git a/cvsps.c b/cvsps.c\nindex 1e64e3c..c22147e 100644\n--- a/cvsps.c\n+++ b/cvsps.c\n@@ -2384,8 +2384,31 @@ void patch_set_add_member(PatchSet * ps,\n     for (next = ps->members.next; next != &ps->members; next = next->next) \n     {\n \tPatchSetMember * m = list_entry(next, PatchSetMember, link);\n-\tif (m->file == psm->file && ps->collision_link.next == NULL) \n-\t\tlist_add(&ps->collision_link, &collisions);\n+\tif (m->file == psm->file) {\n+\t\tint order = compare_rev_strings(psm->post_rev->rev, m->post_rev->rev);\n+\n+\t\t/*\n+\t\t * Same revision too? Add it to the collision list\n+\t\t * if it isn't already.\n+\t\t */\n+\t\tif (!order) {\n+\t\t\tif (ps->collision_link.next == NULL)\n+\t\t\t\tlist_add(&ps->collision_link, &collisions);\n+\t\t\treturn;\n+\t\t}\n+\n+\t\t/*\n+\t\t * If this is an older revision than the one we already have\n+\t\t * in this patchset, just ignore it\n+\t\t */\n+\t\tif (order < 0)\n+\t\t\treturn;\n+\n+\t\t/*\n+\t\t * This is a newer one, remove the old one\n+\t\t */\n+\t\tlist_del(&m->link);\n+\t}\n     }\n \n     psm->ps = ps;\n\ncommit 3d1ebcef6b4f9f6c9064efd64da4dd30d93c3c96\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\nDate:   Wed Mar 22 17:20:20 2006 -0800\n\n    Fix branch ancestor calculation\n    \n    Not having any ancestor at all means that any valid ancestor (even of\n    \"depth 0\") is fine.\n    \n    Signed-off-by: Linus Torvalds <torvalds@osdl.org>\n---\n cvsps.c |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/cvsps.c b/cvsps.c\nindex c22147e..2695a0f 100644\n--- a/cvsps.c\n+++ b/cvsps.c\n@@ -2599,7 +2599,7 @@ static void determine_branch_ancestor(Pa\n \t * note: rev is the pre-commit revision, not the post-commit\n \t */\n \tif (!head_ps->ancestor_branch)\n-\t    d1 = 0;\n+\t    d1 = -1;\n \telse if (strcmp(ps->branch, rev->branch) == 0)\n \t    continue;\n \telse if (strcmp(head_ps->ancestor_branch, \"HEAD\") == 0)\n\n\nuncommitted diff\nAuthor: Linus Torvalds <torvalds@g5.osdl.org>\n\n    Probably totally broken dot counting\n---\n cvsps.c |   13 ++++++++++---\n 1 files changed, 10 insertions(+), 3 deletions(-)\n\ndiff --git a/cvsps.c b/cvsps.c\nindex 2695a0f..2ad1595 100644\n--- a/cvsps.c\n+++ b/cvsps.c\n@@ -2357,9 +2357,16 @@ static int revision_affects_branch(CvsFi\n static int count_dots(const char * p)\n {\n     int dots = 0;\n+    int len = strlen(p);\n \n-    while (*p)\n-\tif (*p++ == '.')\n+    while (len > 2) {\n+\tif (memcmp(p+len-2, \".1\", 2))\n+\t\tbreak;\n+\tlen -= 2;\n+    }\n+\n+    while (len)\n+\tif (p[--len] == '.')\n \t    dots++;\n \n     return dots;\n@@ -2613,7 +2620,7 @@ static void determine_branch_ancestor(Pa\n \t/* HACK: we sometimes pretend to derive from the import branch.  \n \t * just don't do that.  this is the easiest way to prevent... \n \t */\n-\td2 = (strcmp(rev->rev, \"1.1.1.1\") == 0) ? 0 : count_dots(rev->rev);\n+\td2 = count_dots(rev->rev);\n \t\n \tif (d2 > d1)\n \t    head_ps->ancestor_branch = rev->branch;\n"},{"id":"20419","messageId":"Pine.LNX.4.64.0605211844080.3649@g5.osdl.org","threadId":"4219","inReplyTo":"20060520203911.GI6535@nowhere.earth","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T01:45:13Z","receivedAt":"2006-05-22T01:45:13Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 20 May 2006, Yann Dirson wrote:\n> \n> FWIW, I have mentionned a problem that may be the same, under\n> Message-ID <20060107090148.GB32585@nowhere.earth>, that was on January\n> 7th.  Namely, when importing a repository with very large files over\n> pserver or ssh, timeouts can occur and prevent the import from\n> working.  But, as you said, it's not easy to get precise info from the\n> logs :)\n\nFor big repositories, you really shouldn't use pserver or ssh anyway. You \nshould try really really hard to just get a local copy, and do it that \nway. It's going to be tons faster, and will avoid a lot of the problems, \nincluding network timeouts etc.\n\n\t\tLinus\n"},{"id":"20420","messageId":"Pine.LNX.4.64.0605212053590.3697@g5.osdl.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605211209080.3649@g5.osdl.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T03:59:23Z","receivedAt":"2006-05-22T03:59:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 21 May 2006, Linus Torvalds wrote:\n> \n> Ok. It's still converting (that's a big archive), but it has passed the \n> cvsps stage without errors for me, and the conversion so far seems ok. But \n> it has only gotten to \n> \n> \tAuthor: vapier <vapier>  2002-09-23 12:32:42\n> \tChanged GPL to GPL-2 in LICENSE and updated SRC_URI to use mirror:\n> \n> so it has converted only slightly more than the first two years of \n> history in the roughly 30 minutes I've let it run. So it will take several \n> hours.\n\nBtw, trying this import (which got interrupted by a thunderstorm and one \nof our first power failures in a long time - just a few seconds, but \nenough to power off everything but my laptops) it became very obvious that \n\"git cvsimport\" really _really_ should re-pack the archive every once in a \nwhile.\n\nThe old \"repack every month or so\" approach doesn't work that well when \nyou try to import several years of history in a few hours.\n\nNow, you can just repack after the whole thing is done (it will probably \ntake no more than ~15 minutes or so), but it would probably be best if the \nimport script itself decided to repack every once in a while just to avoid \nwasting a lot of diskspace _during_ the import itself.\n\nSo this isn't so much a correctness issue as a \"avoid wasting time and \nspace\" issue, but still..\n\n\t\t\tLinus\n"},{"id":"20421","messageId":"44713BE4.9040505@gentoo.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605212053590.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-22T04:19:48Z","receivedAt":"2006-05-22T04:19:48Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Sun, 21 May 2006, Linus Torvalds wrote:\n>> Ok. It's still converting (that's a big archive), but it has passed the \n>> cvsps stage without errors for me, and the conversion so far seems ok. But \n>> it has only gotten to \n>>\n>> \tAuthor: vapier <vapier>  2002-09-23 12:32:42\n>> \tChanged GPL to GPL-2 in LICENSE and updated SRC_URI to use mirror:\n>>\n>> so it has converted only slightly more than the first two years of \n>> history in the roughly 30 minutes I've let it run. So it will take several \n>> hours.\n> \n> Btw, trying this import (which got interrupted by a thunderstorm and one \n> of our first power failures in a long time - just a few seconds, but \n> enough to power off everything but my laptops) it became very obvious that \n> \"git cvsimport\" really _really_ should re-pack the archive every once in a \n> while.\n\nFortunately the storms haven't been that bad down in Corvallis. cvsps\nalso worked fine for me, but git-cvsimport broke in the middle. The\ncommand I'm using is 'git-cvsimport -P ../gentoo.cvsps -k -d\n/media/scm_comparison -A ~/dev/Authors -v gentoo-x86 | tee cvsimport.log'\n\nHere's the last bits:\n\nFetching gnome-base/gnome-applets/gnome-applets-1.4.0.4-r1.ebuild   v 1.5\nUpdate gnome-base/gnome-applets/gnome-applets-1.4.0.4-r1.ebuild: 947 bytes\nFetching gnome-base/gnome-applets/gnome-applets-1.4.0.4-r2.ebuild   v 1.3\nUpdate gnome-base/gnome-applets/gnome-applets-1.4.0.4-r2.ebuild: 977 bytes\nFetching gnome-base/gnome-applets/gnome-applets-2.0.0-r1.ebuild   v 1.2\nUpdate gnome-base/gnome-applets/gnome-applets-2.0.0-r1.ebuild: 2704 bytes\nFetching gnome-base/gnome-applets/gnome-applets-2.0.0.ebuild   v 1.2\nUpdate gnome-base/gnome-applets/gnome-applets-2.0.0.ebuild: 3031 bytes\nTree ID 4d19a84efce2de9cfb42ac0397e0036bbed2ad65\nParent ID ecb78bbe30369a76e2599d0d17de8fe922dca211\nCommitted patch 14615 (origin 2002-07-16 20:13:15)\nCommit ID 4dd2179e0c1369e07cd268fb5c8b150c3a2a1094\nDelete net-fs/openafs/openafs-1.2.2-r6.ebuild\nDelete net-fs/openafs/files/digest-openafs-1.2.2-r6\nTree ID bfc7320883983655d7d2ea2c6d04f85b45365ce1\nParent ID 4dd2179e0c1369e07cd268fb5c8b150c3a2a1094\nCommitted patch 14616 (origin 2002-07-16 20:15:15)\nCommit ID 7a36de9c4c9b93337ed789ae2341cad3d0991c6d\nUnknown: error  Cannot allocate memory\nFetching profiles/package.mask   v 1.992\ncat: write error: Broken pipe\n\nThanks,\nDonnie\n\n"},{"id":"20424","messageId":"Pine.LNX.4.64.0605212132570.3697@g5.osdl.org","threadId":"4219","inReplyTo":"44713BE4.9040505@gentoo.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T04:50:49Z","receivedAt":"2006-05-22T04:50:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 21 May 2006, Donnie Berkholz wrote:\n> \n> Fortunately the storms haven't been that bad down in Corvallis. cvsps\n> also worked fine for me, but git-cvsimport broke in the middle.\n\nHmm. It's actually possible that it did that for me too - I had put the \ncvsimport in an xterm and forgotten about it, and just assumed that the \npower failure was what broke it. But maybe it had broken down before that \nhappened - I just don't have any logs left ;)\n\n> Here's the last bits:\n> \n> [ snip snip ]\n> Commit ID 7a36de9c4c9b93337ed789ae2341cad3d0991c6d\n> Unknown: error  Cannot allocate memory\n> Fetching profiles/package.mask   v 1.992\n> cat: write error: Broken pipe\n\nHmm. I don't actually know perl, and my original \"cvsimport\" script was \nactually this funny C program that generated a shell script to do the \nimport. That worked fine, and had no memory leaks, but it was a truly \nhacky thing of horrible beauty. Or rather, it _would_ have been that, if \nit had had any beauty to be horrible about. But at least I would have been \nable to debug it.\n\nBut the perl one I can't parse any more. That said, the whole \"Unknown:\" \nprintout seems to come from the subroutine \"_line()\", which just reads a \nline from the cvs server.\n\nDid you do a \"top\" at any time just before this all happened? It _sounds_ \nlike it might actually be a memory leak on the CVS server side, and the \nproblem may (or may not) be due to the optimization that keeps a single \nlong-running CVS server instance for the whole process.\n\nI wouldn't be in the least surprised if that ends up triggering a slow \nleak in CVS itself, and then CVS runs out of memory.\n\nThat would likely have been obvious in any \"top\" output just before the \nfailure.\n\nSmurf, Martin, Dscho.. Any ideas? My old script just ran RCS directly on \nthe files, and had no issues like that. I'll happily admit that my old \nscript generator thing was horrible, but it was a lot easier to debug than \nthe smarter perl script that uses a CVS server connection..\n\n\t\tLinus\n"},{"id":"20425","messageId":"46a038f90605212204m7735d637j50aeb9807eae336a@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605212132570.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T05:04:34Z","receivedAt":"2006-05-22T05:04:34Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/22/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> I wouldn't be in the least surprised if that ends up triggering a slow\n> leak in CVS itself, and then CVS runs out of memory.\n\nI'm dying to try this out myself after work. I don't discard that\ncvsimport might be stuffing data in an array that grows forever. In\nany case you'll hear from me soon.\n\n\n\nmartin\n"},{"id":"20426","messageId":"44714A3D.5060909@gentoo.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605212132570.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-22T05:21:01Z","receivedAt":"2006-05-22T05:21:01Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Linus Torvalds wrote:\n> Did you do a \"top\" at any time just before this all happened? It _sounds_ \n> like it might actually be a memory leak on the CVS server side, and the \n> problem may (or may not) be due to the optimization that keeps a single \n> long-running CVS server instance for the whole process.\n\nNo. =\\ I just started the thing running in a screen session and came\nback a few hours later to find it like that.\n\nThanks,\nDonnie\n\n"},{"id":"20429","messageId":"46a038f90605220042v369e9ff5o3dc7841472171d02@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605212132570.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T07:42:10Z","receivedAt":"2006-05-22T07:42:10Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/22/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Did you do a \"top\" at any time just before this all happened? It _sounds_\n> like it might actually be a memory leak on the CVS server side, and the\n> problem may (or may not) be due to the optimization that keeps a single\n> long-running CVS server instance for the whole process.\n\nRunning a few tests right now. Looks like cvs (Debian/etch 1.12.9-13)\nitself is not leaking any memory. The Perl (Debian/etch\n5.8.7-something and now 5.8.8-4) process OTOH is visibly allocating\nmemory. Starts off at 4MB and gets up to ~17MB by the time it has done\n6K commits.\n\nI am trying to figure out whether the leak is in the script or in the\nPerl implementation, using PadWalk, Devel::Leak and friends. If the\nleak is here, I can't see it (yet).\n\n> I wouldn't be in the least surprised if that ends up triggering a slow\n> leak in CVS itself, and then CVS runs out of memory.\n\nOr a slow leak in Perl? The 5.8.8 release notes do talk about some\nleaks being fixed, but this 5.8.8 isn't making a difference.\n\nWorking on it.\n\n\n\nmartin\n"},{"id":"20436","messageId":"Pine.LNX.4.64.0605220203200.3697@g5.osdl.org","threadId":"4219","inReplyTo":"46a038f90605220042v369e9ff5o3dc7841472171d02@mail.gmail.com","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T09:13:08Z","receivedAt":"2006-05-22T09:13:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 22 May 2006, Martin Langhoff wrote:\n> \n> Or a slow leak in Perl? The 5.8.8 release notes do talk about some\n> leaks being fixed, but this 5.8.8 isn't making a difference.\n> \n> Working on it.\n\nThanks. Looking at what I did convert, that horrid gentoo CVS tree is \ninteresting. The resulting (partial) git history has 93413 commits and \n850,000+ objects total, all in a totally linear history.\n\nAnd that's just up to April 2004, so the full tree is probably a million \nobjects.\n\nThe good news is that git seems to handle that size repo no problem at \nall. The repack did indeed take a long while, but it packed it all down to \na 189MB pack-file (and 20MB pack index).\n\nConsidering that the bzip2'd tar-file of the CVS history was 157MB, and \nthe actual CVS footprint was about 1.6GB, if git stays at under a quarter \ngigabyte for the whole archive once converted (which sounds likely, \ncounting indexing), git would basically cut down the disk usage for a live \nrepo by a factor of 7 or so.\n\n_And_ I can do a \"git log origin > /dev/null\" in about 2.4 seconds. Take \nthat, CVS.\n\n\t\tLinus\n"},{"id":"20445","messageId":"46a038f90605220554y569c11b9p24027772bd2ee79a@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605220203200.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T12:54:15Z","receivedAt":"2006-05-22T12:54:15Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/22/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> On Mon, 22 May 2006, Martin Langhoff wrote:\n> >\n> > Or a slow leak in Perl? The 5.8.8 release notes do talk about some\n> > leaks being fixed, but this 5.8.8 isn't making a difference.\n> >\n> > Working on it.\n>\n> Thanks. Looking at what I did convert, that horrid gentoo CVS tree is\n> interesting. The resulting (partial) git history has 93413 commits and\n> 850,000+ objects total, all in a totally linear history.\n\nOk, so there's 3 patches posted that should help narrow down the\nproblem. There's a new -L <imit> so that Donnie can get his stuff done\nby running it in a while(true) loop. Not proud of it, but hey.\n\nAnd there are two patches that I suspect may fix the leak. After\napplying them, the cvsimport process grows up to ~13MB and then tapers\noff, at least as far as my patience has gotten me. It's late on this\nside of the globe so I'll look at the results tomorrow morning.\n\n(BTW, I typo-ed Linus' address in the git-send-email invocation. Will\nresend to him separately)\n\nI'll also prep a patch as Linus suggests to do auto-repacking while\nthe import runs so we don't eat up the harddisk.\n\n> git would basically cut down the disk usage for a live\n> repo by a factor of 7 or so.\n>\n> _And_ I can do a \"git log origin > /dev/null\" in about 2.4 seconds. Take\n> that, CVS.\n\nHeh. Faster Gitticat, Kill Kill Kill!\n\n\n\n\nmartin\n"},{"id":"20458","messageId":"Pine.LNX.4.64.0605221013020.3697@g5.osdl.org","threadId":"4219","inReplyTo":"46a038f90605220554y569c11b9p24027772bd2ee79a@mail.gmail.com","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T17:27:01Z","receivedAt":"2006-05-22T17:27:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 23 May 2006, Martin Langhoff wrote:\n> \n> And there are two patches that I suspect may fix the leak. After\n> applying them, the cvsimport process grows up to ~13MB and then tapers\n> off, at least as far as my patience has gotten me. It's late on this\n> side of the globe so I'll look at the results tomorrow morning.\n\nOk, initial results are promising. git-cvsimport appears to be still \nslowly growing, but it's at 40M (ie pretty tiny, considering that cvsps \ngrew to 800+MB on this archive) and growth seems to actually be slowing.\n\nMy conversion is only up to September 2002, but if it doesn't suddenly hit \nsome huge growth spurt, I wouldn't expect it to run out of memory. The CVS \nserver process itself is tiny, and doesn't seem to grow at all.\n\nAs to packing, it doing something like\n\n\twhile :\n\tdo\n\t\tsleep 30\n\n\t\t#\n\t\t# repack roughly every 25600 objects\n\t\t#\n\t\tn=$(ls .git/objects/00 2> /dev/null | wc -l)\n\t\tif [ $n -gt 100 ]; then\n\t\t\tgit repack -a\n\t\t\t#\n\t\t\t# Stupid sleep to make sure that nobody is still\n\t\t\t# using any unpacked objects after the pack got\n\t\t\t# generated\n\t\t\t#\n\t\t\tsleep 10\n\t\t\tgit prune-packed\n\t\tfi\n\tdone\n\nor similar (the above is totally untested - I've just done it by hand a \nfew times) should work. It's perfectly ok to repack the archive even while \nthe cvsimport script is adding more data and changing it.\n\n\t\tLinus\n"},{"id":"20459","messageId":"e4stna$o1g$1@sea.gmane.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221013020.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-22T17:51:42Z","receivedAt":"2006-05-22T17:51:42Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n\n>                       git repack -a\n>                       #\n>                       # Stupid sleep to make sure that nobody is still\n>                       # using any unpacked objects after the pack got\n>                       # generated\n>                       #\n>                       sleep 10\n>                       git prune-packed\n\nIs it really necessary (on Linux at least)? Git boast it's atomicity...\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"20460","messageId":"Pine.LNX.4.64.0605221055270.3697@g5.osdl.org","threadId":"4219","inReplyTo":"e4stna$o1g$1@sea.gmane.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T18:03:05Z","receivedAt":"2006-05-22T18:03:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 22 May 2006, Jakub Narebski wrote:\n>\n> Linus Torvalds wrote:\n> \n> >                       git repack -a\n> >                       #\n> >                       # Stupid sleep to make sure that nobody is still\n> >                       # using any unpacked objects after the pack got\n> >                       # generated\n> >                       #\n> >                       sleep 10\n> >                       git prune-packed\n> \n> Is it really necessary (on Linux at least)? Git boast it's atomicity...\n\nI don't think it's necessary in practice.\n\nBut people _should_ realize that removing objects is very very special. \nWhether it's done by \"git prune-packed\" or \"git prune\", that's a very \ndangerous operations. \"git prune\" a lot more so than \"git prune-packed\", \nof course (in fact, you should _never_ run \"git prune\" on a repository \nthat is active - you _will_ corrupt it)-\n\nDoing \"git prune-packed\" _should_ be mostly safe on UNIX, since the \nobjects all exist in packs, and anybody who already opened an object will \nkeep the fd open, and not even notice that the name is gone. However, \nthere is at least one race:\n\n\tobject lookup\t\t\t\"git repack -a -d\"\n\t=============\t\t\t==================\n\n - a process does its object\n   database setup. No new pack-file\n   yet.\n\n\t\t\t\t\t - mv tmp-packfile active-packfile\n\n\t\t\t\t\t - git prune-packed\n\n - the process looks up the object,\n   and doesn't look in the pack-file\n   because it didn't see the pack-file.\n\n   So it tries to look up an object,\n   fails, and errors out.\n\n   It's not a fatal error (just re-try)\n   but it could break something like a\n   cvsimport\n\nNow, in PRACTICE, I doubt you'd ever hit this. But the fact is, pruning \nyour repository (whether prune-packed or a full prune) is _the_ special \noperation. It's something that removes a filesystem representation of an \nobject that is otherwise immutable.\n\n\t\tLinus\n"},{"id":"20461","messageId":"E1FiFgL-0003m6-Eb@moooo.ath.cx","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221055270.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Matthias Lederhofer","fromEmail":"matled@gmx.net","sentAt":"2006-05-22T19:03:09Z","receivedAt":"2006-05-22T19:03:09Z","isPatch":false,"sender":{"key":"matled@gmx.net","avatar":null},"body":"> But people _should_ realize that removing objects is very very special. \n\nJust a similar question: is there any reason not tu run git\nrepack/prune-packed as cron job? I would think of something like this\nfor every night:\n\n- git prune-packed (remove objects packed last time)\n- check how many objects git-count-objects counts, if it are not enough\n  abort\n- git repack\n\ngit repack -a -d is probably a bad idea, I guess, because a program\ncould try to open them after they were deleted.  Is there any way to\ndelete unnecessary packs (those which would repack -a -d delete)?\nMaking it possible to do a git repack -a and delete those packs the\nnext night?\n"},{"id":"20463","messageId":"7vodxpc1wa.fsf@assigned-by-dhcp.cox.net","threadId":"4219","inReplyTo":"E1FiFgL-0003m6-Eb@moooo.ath.cx","subject":"Re: irc usage..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-22T19:09:25Z","receivedAt":"2006-05-22T19:09:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Matthias Lederhofer <matled@gmx.net> writes:\n\n> ...  Is there any way to\n> delete unnecessary packs (those which would repack -a -d delete)?\n> Making it possible to do a git repack -a and delete those packs the\n> next night?\n\npack-redundant is supposed to figure it out, but I have never\nused it myself so your mileage may vary.\n"},{"id":"20462","messageId":"44720C66.6040304@gentoo.org","threadId":"4219","inReplyTo":"46a038f90605220554y569c11b9p24027772bd2ee79a@mail.gmail.com","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-22T19:09:26Z","receivedAt":"2006-05-22T19:09:26Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 5/22/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> On Mon, 22 May 2006, Martin Langhoff wrote:\n>> >\n>> > Or a slow leak in Perl? The 5.8.8 release notes do talk about some\n>> > leaks being fixed, but this 5.8.8 isn't making a difference.\n>> >\n>> > Working on it.\n>>\n>> Thanks. Looking at what I did convert, that horrid gentoo CVS tree is\n>> interesting. The resulting (partial) git history has 93413 commits and\n>> 850,000+ objects total, all in a totally linear history.\n> \n> Ok, so there's 3 patches posted that should help narrow down the\n> problem. There's a new -L <imit> so that Donnie can get his stuff done\n> by running it in a while(true) loop. Not proud of it, but hey.\n> \n> And there are two patches that I suspect may fix the leak. After\n> applying them, the cvsimport process grows up to ~13MB and then tapers\n> off, at least as far as my patience has gotten me. It's late on this\n> side of the globe so I'll look at the results tomorrow morning.\n\nOK, I started a new run without -L, and I'm watching it in top right\nnow. The cvsimport seems to be doing alright, but the cvs server process\nsucks about another megabyte of virtual every 4-5 seconds. This is a bit\nconcerning since I don't have any swap. Shortly after it hit 670M, I got\n\"Cannot allocate memory\" again. I've got a gig of RAM, and around 300M\nwas resident in various processes at the time.\n\nSo it seems the problem is in cvs itself. I will try another run with -L\nnow.\n\nThanks,\nDonnie\n\n"},{"id":"20468","messageId":"Pine.LNX.4.64.0605221234430.3697@g5.osdl.org","threadId":"4219","inReplyTo":"44720C66.6040304@gentoo.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T19:38:32Z","receivedAt":"2006-05-22T19:38:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 22 May 2006, Donnie Berkholz wrote:\n> \n> OK, I started a new run without -L, and I'm watching it in top right\n> now. The cvsimport seems to be doing alright, but the cvs server process\n> sucks about another megabyte of virtual every 4-5 seconds. This is a bit\n> concerning since I don't have any swap. Shortly after it hit 670M, I got\n> \"Cannot allocate memory\" again. I've got a gig of RAM, and around 300M\n> was resident in various processes at the time.\n\nHmm. My cvs server doesn't really grow at all. It's at 13M RSS.\n\nWhat version of cvs are you running?\n\n\t[torvalds@g5 ~]$ cvs --version\n\n\tConcurrent Versions System (CVS) 1.11.21 (client/server)\n\nmaybe that matters.\n\n(but my import is only up to Jun 22, 2003 so far).\n\n\t\tLinus\n"},{"id":"20469","messageId":"46a038f90605221241x58ffa2a4o26159d38d86a8092@mail.gmail.com","threadId":"4219","inReplyTo":"44720C66.6040304@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T19:41:38Z","receivedAt":"2006-05-22T19:41:38Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/23/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> So it seems the problem is in cvs itself. I will try another run with -L\n> now.\n\nWhat version of cvs are you using? Perhaps trying a different one?\n\nThe dev machine where I am running the import is a slug! It's still\nworking on it, only gotten to 7700 commits, with the cvsimport process\nstable at 28MB RAM and cvs stable at 4MB.\n\ncheers,\n\n\nmartin\n"},{"id":"20470","messageId":"46a038f90605221246l6df6662dsb79d524ba2f0780d@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221013020.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T19:46:23Z","receivedAt":"2006-05-22T19:46:23Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/23/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Ok, initial results are promising. git-cvsimport appears to be still\n> slowly growing, but it's at 40M (ie pretty tiny, considering that cvsps\n> grew to 800+MB on this archive) and growth seems to actually be slowing.\n\nThat's great news. The cvs archive seems to have large commits every\nonce in a while, so I suspect the residual memory growth may be\nrelated to those. Or to a smaller leak I haven't nailed.\n\nMy test box is bloody slow it seems. I'll try and get hold of a faster\nmachine to run this if I can.\n\n> As to packing, it doing something like\n\nGiven that we are running batch, it is safe and simple to stop the\nimport, repack, prune-packed, and keep going. Don't think we'll win\nany races by running it in parallel ;-)\n\ncheers,\n\n\nmartin\n"},{"id":"20471","messageId":"447215D4.5020403@gentoo.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221234430.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-22T19:49:40Z","receivedAt":"2006-05-22T19:49:40Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Linus Torvalds wrote:\n> Hmm. My cvs server doesn't really grow at all. It's at 13M RSS.\n\nYeah, that's the thing. RSS stayed about the same (according to top),\nbut virtual just kept growing.\n\n> What version of cvs are you running?\n> \n> \t[torvalds@g5 ~]$ cvs --version\n> \n> \tConcurrent Versions System (CVS) 1.11.21 (client/server)\n\nConcurrent Versions System (CVS) 1.12.12 (client/server)\n\nLooks like there's a .13 out but the zlib interaction is badly broken\n(-z >=1) so my system didn't get upgraded. I'll try it anyway after the\n-L run finishes.\n\nThanks,\nDonnie\n\n"},{"id":"20472","messageId":"Pine.LNX.4.64.0605221256090.3697@g5.osdl.org","threadId":"4219","inReplyTo":"46a038f90605221241x58ffa2a4o26159d38d86a8092@mail.gmail.com","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T20:11:18Z","receivedAt":"2006-05-22T20:11:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 23 May 2006, Martin Langhoff wrote:\n> \n> The dev machine where I am running the import is a slug! It's still\n> working on it, only gotten to 7700 commits, with the cvsimport process\n> stable at 28MB RAM and cvs stable at 4MB.\n\nI have to say, that cvsimport script really does do horrible things. It's \nbasically a fork/exec/exit benchmark, as far as I can tell. Running \noprofile on the thing, the top offenders are (ignore the 45% idle thing: \nit's just because this was run on a dual-cpu system, so since it's almost \ncompletely single-threaded you get ~50% idle by default).\n\n\t3117654  45.8708  vmlinux                  vmlinux                  .power4_idle\n\t802313   11.8046  vmlinux                  vmlinux                  .unmap_vmas\n\t632913    9.3122  vmlinux                  vmlinux                  .copy_page_range\n\t150359    2.2123  vmlinux                  vmlinux                  .release_pages\n\t131330    1.9323  vmlinux                  vmlinux                  .vm_normal_page\n\t117836    1.7337  libperl.so               libperl.so               (no symbols)\n\t74098     1.0902  libgklayout.so           libgklayout.so           (no symbols)\n\t54680     0.8045  vmlinux                  vmlinux                  .free_pages_and_swap_cache\n\t54300     0.7989  libfb.so                 libfb.so                 (no symbols)\n\t49052     0.7217  vmlinux                  vmlinux                  .copy_4K_page\n\t46559     0.6850  libc-2.4.so              libc-2.4.so              getc\n\t42677     0.6279  vmlinux                  vmlinux                  .page_remove_rmap\n\t41133     0.6052  libc-2.4.so              libc-2.4.so              ferror\n\t..\n\nthose kernel functions are all about process create/exit, and COW faulting \nafter the fork.\n\nNow, this is on ppc, so process creation is likely slower (idiotic PPC VM \npage table hashes), but Linux is actually very good at doing this, and the \nfact that process create/exit is so high is a very big sign that the \nscript just ends up executing a _ton_ of small simple processes that do \nalmost nothing.\n\nI wonder why those \"git-update-index\" calls seem to be (assuming I read \nthe perl correctly) done only a few files at a time. We can do a hundreds \nin one go, but it seems to want to do just ten files or something at the \nsame time. Although since most commits should hopefully just modify a \ncouple of files, that probably isn't a big deal.\n\nThat thing would probably be an order of magnitude faster if written to \nuse the git library interfaces directly. Of course, the CVS part is \nprobably a big overhead, so it might not help much (I would not be \nsurprised at all if a number of the fork/exec/exit things are due to the \nCVS server starting RCS or something, not due to git-cvsimport itself)\n\n\t\tLinus\n"},{"id":"20473","messageId":"44721C26.5000503@gentoo.org","threadId":"4219","inReplyTo":"44720C66.6040304@gentoo.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-22T20:16:38Z","receivedAt":"2006-05-22T20:16:38Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Donnie Berkholz wrote:\n> OK, I started a new run without -L, and I'm watching it in top right\n> now.\n\nTried a run with -L 1024 and it broke in just a couple of minutes:\n\nFetching\nsys-kernel/linux/files/2.4.0.8/linux-2.4.0-ac8-reiserfs-3.6.25-nfs.diff.gz\n  v 1.1\nNew\nsys-kernel/linux/files/2.4.0.8/linux-2.4.0-ac8-reiserfs-3.6.25-nfs.diff.gz:\n6367 bytes\nTree ID 457f629df10e70a5ef430f431eca27ed02a83d46\nParent ID 0541d8b54a02df3be50d529497236556c6862a4c\nCommitted patch 1024 (origin 2001-01-13 00:29:39)\nCommit ID ba9d995d12a37502a851e198b67e141623f79544\nDONE; creating master branch\ncat: write error: Broken pipe\n\nThanks,\nDonnie\n\n"},{"id":"20474","messageId":"Pine.LNX.4.64.0605221312380.3697@g5.osdl.org","threadId":"4219","inReplyTo":"447215D4.5020403@gentoo.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T20:20:29Z","receivedAt":"2006-05-22T20:20:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 22 May 2006, Donnie Berkholz wrote:\n>\n> Linus Torvalds wrote:\n> > Hmm. My cvs server doesn't really grow at all. It's at 13M RSS.\n> \n> Yeah, that's the thing. RSS stayed about the same (according to top),\n> but virtual just kept growing.\n\nNot for me. The virtual size is certainly bigger than RSS, but not by a \nhuge amount. So this might be a regression in CVS, since you seem to have \na newer version than I do.\n\nThe latest stable CVS release is 1.11.21, I think: you seem to be running \nthe \"development\" version (1.12.x).\n\n\t\t\tLinus\n"},{"id":"20475","messageId":"Pine.LNX.4.64.0605221328210.3697@g5.osdl.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221256090.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T20:33:01Z","receivedAt":"2006-05-22T20:33:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 22 May 2006, Linus Torvalds wrote:\n> \n> Of course, the CVS part is probably a big overhead, so it might not help \n> much (I would not be surprised at all if a number of the fork/exec/exit \n> things are due to the CVS server starting RCS or something, not due to \n> git-cvsimport itself)\n\nAhh. stracing the CVS server seems to imply that it forks off a subprocess \nfor every command. It doesn't actually execute any external program, but \njust does a fork + muck around in the ,v files + exit.\n\nMaybe one of the changes in the 1.12.x versions is to not do that, which \nmight explain why Donnie seems to see much better performance, but also \nsees all the memory leakage?\n\n\t\tLinus\n"},{"id":"20482","messageId":"20060522214128.GE16677@kiste.smurf.noris.de","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221256090.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Matthias Urlichs","fromEmail":"smurf@smurf.noris.de","sentAt":"2006-05-22T21:41:28Z","receivedAt":"2006-05-22T21:41:28Z","isPatch":false,"sender":{"key":"matthias@urlichs.de","avatar":"https://gravatar.com/avatar/2708905af227313eba6f2b2ae0f7d0259b5ac5d71baef58fe5a13c699ce0bbf0?d=mp&s=160"},"body":"Hi,\n\nLinus Torvalds:\n> I wonder why those \"git-update-index\" calls seem to be (assuming I read \n> the perl correctly) done only a few files at a time. We can do a hundreds \n> in one go, but it seems to want to do just ten files or something at the \n> same time.\n\nNo, fifty.\n\nI simply was too lazy to count the actual filenames' lengths. ;-)\n\n> That thing would probably be an order of magnitude faster if written to \n> use the git library interfaces directly. Of course, the CVS part is \n> probably a big overhead, so it might not help much \n\nThe beast *was* mainly written to do this remotely...\n\n-- \nMatthias Urlichs   |   {M:U} IT Design @ m-u-it.de   |  smurf@smurf.noris.de\nDisclaimer: The quote was selected randomly. Really. | http://smurf.noris.de\n - -\nThe worst form of inequality is to try to make unequal things equal.\n\t\t\t\t\t-- Aristotle\n"},{"id":"20478","messageId":"447231C4.2030508@gentoo.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221312380.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-22T21:48:52Z","receivedAt":"2006-05-22T21:48:52Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Linus Torvalds wrote:\n> The latest stable CVS release is 1.11.21, I think: you seem to be running \n> the \"development\" version (1.12.x).\n\nBacked down to the 1.11 series, things seem to be going fine so far.\n\nThanks,\nDonnie\n\n"},{"id":"20483","messageId":"Pine.LNX.4.64.0605221516500.3697@g5.osdl.org","threadId":"4219","inReplyTo":"20060522214128.GE16677@kiste.smurf.noris.de","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T22:18:35Z","receivedAt":"2006-05-22T22:18:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 22 May 2006, Matthias Urlichs wrote:\n> \n> The beast *was* mainly written to do this remotely...\n\nI don't think the remote usability is valid, except for some really small \nrepositories. The fact that it takes hours even when the CVS server is \nlocal doesn't bode well for doing it remotely for any but the most trivial \nthings.\n\nI really think it would be better to have local use be the optimized case, \nwith remote being the \"it's _possible_\" case.\n\n\t\tLinus\n"},{"id":"20485","messageId":"7v8xotadm3.fsf@assigned-by-dhcp.cox.net","threadId":"4219","inReplyTo":"20060522214128.GE16677@kiste.smurf.noris.de","subject":"Re: irc usage..","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-22T22:39:16Z","receivedAt":"2006-05-22T22:39:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Matthias Urlichs <smurf@smurf.noris.de> writes:\n\n> Hi,\n>\n> Linus Torvalds:\n>> I wonder why those \"git-update-index\" calls seem to be (assuming I read \n>> the perl correctly) done only a few files at a time. We can do a hundreds \n>> in one go, but it seems to want to do just ten files or something at the \n>> same time.\n>\n> No, fifty.\n>\n> I simply was too lazy to count the actual filenames' lengths. ;-)\n\nI think cvsimport predates that option, but these days that loop\ncan be optimized by feeding --index-info from standard input.\n"},{"id":"20491","messageId":"46a038f90605221615j59583bcdqf128bab31603148e@mail.gmail.com","threadId":"4219","inReplyTo":"7v8xotadm3.fsf@assigned-by-dhcp.cox.net","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T23:15:07Z","receivedAt":"2006-05-22T23:15:07Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/23/06, Junio C Hamano <junkio@cox.net> wrote:\n> > I simply was too lazy to count the actual filenames' lengths. ;-)\n>\n> I think cvsimport predates that option, but these days that loop\n> can be optimized by feeding --index-info from standard input.\n\nOh, yep, that'd be a good addition. I think we can also cut down on\nthe number of fork+exec calls (as Linus points out they are killing\nus) by caching some data we should already have that we are repeatedly\nasking from git-ref-parse.\n\nOther TODOs from my reading of the code last night...\n\n - Switch from line-oriented reads to block reads when fetching files\nfrom CVS. This gentoo has repo has some large binary blobs in it and\nwe end up slurping them into memory.\n\n - Stop abusing globals in commit() -- pass the commit data as parameters.\n\n - Further profiling? Whatever we are doing, we aren't doing it fast :(\n\nWill be trying to do those things in the next few days, don't mind if\nsomeone jumps in as well.\n\n\n\nmartin\n"},{"id":"20494","messageId":"46a038f90605221623g25325e71hf3faf0a6a6ca628a@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221516500.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T23:23:06Z","receivedAt":"2006-05-22T23:23:06Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/23/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> I don't think the remote usability is valid, except for some really small\n> repositories. The fact that it takes hours even when the CVS server is\n> local doesn't bode well for doing it remotely for any but the most trivial\n> things.\n\nI really don't think that using the local cvs binary is a problem at\nall. In my experience, the thing is fairly fast and optimized when you\nask it to perform file-oriented questions and that's all we do,\nreally.\n\nIf you want to try it, you'll see that local checkouts of large trees\n(like this gentoo one) are fairly fast. Not as fast as GIT itself, but\ngood enough. I think Donnie has hit a bug with a bad version of cvs,\nbut other than that, my experience with it is that it is fairly well\nbehaved -- even if the tool is bad, ubiquity has lead to resiliency\nover the years.\n\n> I really think it would be better to have local use be the optimized case,\n> with remote being the \"it's _possible_\" case.\n\nAgreed, but I think we won't see much benefit in direct parsing. And\nwe'll have to take the hit of double-implementation.\n\nIn any case, we have it already -- parsecvs does it quite well (modulo\nmemory leaks!) and I've used it several times in conjunction with\ncvsimport. Just perform the initial import with parsecvs and then\n'track' the remote project with cvsimport.\n\nThe problem is that they lead to slightly different trees. So their\noutput is not consistent, and I don't think that'll be easy to fix.\n\ncheers,\n\n\nmartin\n"},{"id":"20496","messageId":"46a038f90605221629w4e5b3654o5bb494843cb1d38a@mail.gmail.com","threadId":"4219","inReplyTo":"46a038f90605221623g25325e71hf3faf0a6a6ca628a@mail.gmail.com","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-22T23:29:32Z","receivedAt":"2006-05-22T23:29:32Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/23/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> The problem is that they lead to slightly different trees.\n\nSorry! s/trees/histories/ there. The trees are (or should!) be the\nsame, and tree differences should be addressed as bugs. Differences in\nhow history is parsed are unavoidable right now.\n\nmartin\n"},{"id":"20497","messageId":"Pine.LNX.4.64.0605221629080.3697@g5.osdl.org","threadId":"4219","inReplyTo":"46a038f90605221623g25325e71hf3faf0a6a6ca628a@mail.gmail.com","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-22T23:33:06Z","receivedAt":"2006-05-22T23:33:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 23 May 2006, Martin Langhoff wrote:\n> \n> I really don't think that using the local cvs binary is a problem at\n> all. In my experience, the thing is fairly fast and optimized when you\n> ask it to perform file-oriented questions and that's all we do,\n> really.\n\nFair enough. My worry was mainly that the cvs server was doing something \nstupid, but I suspect most of the fork/exec's are probably from the \ncvsimport perl script itself.\n\n> In any case, we have it already -- parsecvs does it quite well (modulo\n> memory leaks!) and I've used it several times in conjunction with\n> cvsimport. Just perform the initial import with parsecvs and then\n> 'track' the remote project with cvsimport.\n\nI didn't get parsecvs working when I tried it a long time ago, and Donnie \nreported that it ran out of memory, so I didn't even really consider it. \nI'd love for it to work well, and it may be reasonable to do really big \nimports on multi-gigabyte 64-bit machines (after all, they aren't _hard_ \nto find any more, and you only need to do it once).\n\nThat said, it still seems pretty stupid to require that much memory just \nto import from CVS.\n\n\t\tLinus\n"},{"id":"20509","messageId":"20060523065232.GA6180@coredump.intra.peff.net","threadId":"4219","inReplyTo":"46a038f90605221615j59583bcdqf128bab31603148e@mail.gmail.com","subject":"Re: irc usage..","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T06:52:32Z","receivedAt":"2006-05-23T06:52:32Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, May 23, 2006 at 11:15:07AM +1200, Martin Langhoff wrote:\n\n> >I think cvsimport predates that option, but these days that loop\n> >can be optimized by feeding --index-info from standard input.\n> Oh, yep, that'd be a good addition. I think we can also cut down on\n\nThis patch is relatively simple, and I'll post it in a moment.\n\nI also made a few other cleanups to commit() which apply on top of that;\nI'll post it also.\n\n> - Stop abusing globals in commit() -- pass the commit data as parameters.\n\nSome of the globals actually get modified in commit() (e.g., @old and\n@new get cleared).  So we need to either pass them in as references or\nremember to do that cleanup each time it is called (which is really only\ntwice, I think).\n\n> Will be trying to do those things in the next few days, don't mind if\n> someone jumps in as well.\n\nI can look at the line/block CVS file slurping, but not tonight.\n\n-Peff\n"},{"id":"20510","messageId":"20060523065810.GB6180@coredump.intra.peff.net","threadId":"4219","inReplyTo":"20060523065232.GA6180@coredump.intra.peff.net","subject":"Re: irc usage..","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T06:58:10Z","receivedAt":"2006-05-23T06:58:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":">From nobody Mon Sep 17 00:00:00 2001\nFrom: Jeff King <peff@peff.net>\nDate: Tue, 23 May 2006 01:16:07 -0400\nSubject: [PATCH 1/2] cvsimport: use git-update-index --index-info\n\nThis should reduce the number of git-update-index forks required per\ncommit. We now do adds/removes in one call, and we are no longer forced to\ndeal with argv limitations.\n\n---\n\ncb6452bbfda9c52ad8dbeaac6e3440ae61099a05\n git-cvsimport.perl |   36 +++++++++++++-----------------------\n 1 files changed, 13 insertions(+), 23 deletions(-)\n\ncb6452bbfda9c52ad8dbeaac6e3440ae61099a05\ndiff --git a/git-cvsimport.perl b/git-cvsimport.perl\nindex d257e66..4efb0a5 100755\n--- a/git-cvsimport.perl\n+++ b/git-cvsimport.perl\n@@ -565,29 +565,19 @@ my($patchset,$date,$author_name,$author_\n my(@old,@new,@skipped);\n sub commit {\n \tmy $pid;\n-\twhile(@old) {\n-\t\tmy @o2;\n-\t\tif(@old > 55) {\n-\t\t\t@o2 = splice(@old,0,50);\n-\t\t} else {\n-\t\t\t@o2 = @old;\n-\t\t\t@old = ();\n-\t\t}\n-\t\tsystem(\"git-update-index\",\"--force-remove\",\"--\",@o2);\n-\t\tdie \"Cannot remove files: $?\\n\" if $?;\n-\t}\n-\twhile(@new) {\n-\t\tmy @n2;\n-\t\tif(@new > 12) {\n-\t\t\t@n2 = splice(@new,0,10);\n-\t\t} else {\n-\t\t\t@n2 = @new;\n-\t\t\t@new = ();\n-\t\t}\n-\t\tsystem(\"git-update-index\",\"--add\",\n-\t\t\t(map { ('--cacheinfo', @$_) } @n2));\n-\t\tdie \"Cannot add files: $?\\n\" if $?;\n-\t}\n+\n+      \topen(my $fh, '|-', qw(git-update-index --index-info))\n+\t\tor die \"unable to open git-update-index: $!\";\n+\tprint $fh \n+\t\t(map { \"0 0000000000000000000000000000000000000000\\t$_\\n\" }\n+\t\t\t@old),\n+\t\t(map { '100' . sprintf('%o', $_->[0]) . \" $_->[1]\\t$_->[2]\\n\" }\n+\t\t\t@new)\n+\t\tor die \"unable to write to git-update-index: $!\";\n+\tclose $fh\n+\t\tor die \"unable to write to git-update-index: $!\";\n+\t$? and die \"git-update-index reported error: $?\";\n+\t@old = @new = ();\n \n \t$pid = open(C,\"-|\");\n \tdie \"Cannot fork: $!\" unless defined $pid;\n-- \n1.3.3.gcb64-dirty\n"},{"id":"20511","messageId":"20060523070007.GC6180@coredump.intra.peff.net","threadId":"4219","inReplyTo":"20060523065232.GA6180@coredump.intra.peff.net","subject":"[PATCH 2/2] cvsimport: cleanup commit function","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T07:00:07Z","receivedAt":"2006-05-23T07:00:07Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This change attempts to clean up the commit function to make it a bit\neasier to read (or at least the first half of it). It also improves\nrobustness and performance. Specifically:\n  - report get_headref errors on opening ref unless the error is ENOENT\n  - use regex to check for sha1 instead of length\n  - use lexically scoped filehandles which get cleaned up automagically\n  - check for error on both 'print' and 'close' (since output is buffered)\n  - avoid \"fork, do some perl, then exec\" in commit(). It's not necessary,\n    and we probably end up COW'ing parts of the perl process. Plus the code\n    is much smaller because we can use open2()\n  - avoid calling strftime over and over (mainly a readability cleanup)\n\n---\n\nI know this patch is quite large. I can try to split it if you want, but\nI suspect it's not worth the effort (either you like refactoring or you\ndon't :) ).\n\n9dc9f05ab5e1cbd8765238e7b1da0addd6f4296a\n git-cvsimport.perl |  150 ++++++++++++++++++++++------------------------------\n 1 files changed, 64 insertions(+), 86 deletions(-)\n\n9dc9f05ab5e1cbd8765238e7b1da0addd6f4296a\ndiff --git a/git-cvsimport.perl b/git-cvsimport.perl\nindex 4efb0a5..f8feb52 100755\n--- a/git-cvsimport.perl\n+++ b/git-cvsimport.perl\n@@ -23,7 +23,7 @@ use File::Basename qw(basename dirname);\n use Time::Local;\n use IO::Socket;\n use IO::Pipe;\n-use POSIX qw(strftime dup2);\n+use POSIX qw(strftime dup2 :errno_h);\n use IPC::Open2;\n \n $SIG{'PIPE'}=\"IGNORE\";\n@@ -429,22 +429,25 @@ sub getwd() {\n \treturn $pwd;\n }\n \n+sub is_sha1 {\n+\tmy $s = shift;\n+\treturn $s =~ /^[a-zA-Z0-9]{40}$/;\n+}\n \n-sub get_headref($$) {\n+sub get_headref ($$) {\n     my $name    = shift;\n     my $git_dir = shift; \n-    my $sha;\n     \n-    if (open(C,\"$git_dir/refs/heads/$name\")) {\n-\tchomp($sha = <C>);\n-\tclose(C);\n-\tlength($sha) == 40\n-\t    or die \"Cannot get head id for $name ($sha): $!\\n\";\n+    my $f = \"$git_dir/refs/heads/$name\";\n+    if(open(my $fh, $f)) {\n+      \t    chomp(my $r = <$fh>);\n+\t    is_sha1($r) or die \"Cannot get head id for $name ($r): $!\";\n+\t    return $r;\n     }\n-    return $sha;\n+    die \"unable to open $f: $!\" unless $! == POSIX::ENOENT;\n+    return undef;\n }\n \n-\n -d $git_tree\n \tor mkdir($git_tree,0777)\n \tor die \"Could not create $git_tree: $!\";\n@@ -561,90 +564,67 @@ #---------------------\n \n my $state = 0;\n \n-my($patchset,$date,$author_name,$author_email,$branch,$ancestor,$tag,$logmsg);\n-my(@old,@new,@skipped);\n-sub commit {\n-\tmy $pid;\n-\n+sub update_index (\\@\\@) {\n+\tmy $old = shift;\n+\tmy $new = shift;\n       \topen(my $fh, '|-', qw(git-update-index --index-info))\n \t\tor die \"unable to open git-update-index: $!\";\n \tprint $fh \n \t\t(map { \"0 0000000000000000000000000000000000000000\\t$_\\n\" }\n-\t\t\t@old),\n+\t\t\t@$old),\n \t\t(map { '100' . sprintf('%o', $_->[0]) . \" $_->[1]\\t$_->[2]\\n\" }\n-\t\t\t@new)\n+\t\t\t@$new)\n \t\tor die \"unable to write to git-update-index: $!\";\n \tclose $fh\n \t\tor die \"unable to write to git-update-index: $!\";\n \t$? and die \"git-update-index reported error: $?\";\n-\t@old = @new = ();\n+}\n \n-\t$pid = open(C,\"-|\");\n-\tdie \"Cannot fork: $!\" unless defined $pid;\n-\tunless($pid) {\n-\t\texec(\"git-write-tree\");\n-\t\tdie \"Cannot exec git-write-tree: $!\\n\";\n-\t}\n-\tchomp(my $tree = <C>);\n-\tlength($tree) == 40\n-\t\tor die \"Cannot get tree id ($tree): $!\\n\";\n-\tclose(C)\n+sub write_tree () {\n+\topen(my $fh, '-|', qw(git-write-tree))\n+\t\tor die \"unable to open git-write-tree: $!\";\n+\tchomp(my $tree = <$fh>);\n+\tis_sha1($tree)\n+\t\tor die \"Cannot get tree id ($tree): $!\";\n+\tclose($fh)\n \t\tor die \"Error running git-write-tree: $?\\n\";\n \tprint \"Tree ID $tree\\n\" if $opt_v;\n+\treturn $tree;\n+}\n \n-\tmy $parent = \"\";\n-\tif(open(C,\"$git_dir/refs/heads/$last_branch\")) {\n-\t\tchomp($parent = <C>);\n-\t\tclose(C);\n-\t\tlength($parent) == 40\n-\t\t\tor die \"Cannot get parent id ($parent): $!\\n\";\n-\t\tprint \"Parent ID $parent\\n\" if $opt_v;\n-\t}\n-\n-\tmy $pr = IO::Pipe->new() or die \"Cannot open pipe: $!\\n\";\n-\tmy $pw = IO::Pipe->new() or die \"Cannot open pipe: $!\\n\";\n-\t$pid = fork();\n-\tdie \"Fork: $!\\n\" unless defined $pid;\n-\tunless($pid) {\n-\t\t$pr->writer();\n-\t\t$pw->reader();\n-\t\topen(OUT,\">&STDOUT\");\n-\t\tdup2($pw->fileno(),0);\n-\t\tdup2($pr->fileno(),1);\n-\t\t$pr->close();\n-\t\t$pw->close();\n-\n-\t\tmy @par = ();\n-\t\t@par = (\"-p\",$parent) if $parent;\n-\n-\t\t# loose detection of merges\n-\t\t# based on the commit msg\n-\t\tforeach my $rx (@mergerx) {\n-\t\t\tif ($logmsg =~ $rx) {\n-\t\t\t\tmy $mparent = $1;\n-\t\t\t\tif ($mparent eq 'HEAD') { $mparent = $opt_o };\n-\t\t\t\tif ( -e \"$git_dir/refs/heads/$mparent\") {\n-\t\t\t\t\t$mparent = get_headref($mparent, $git_dir);\n-\t\t\t\t\tpush @par, '-p', $mparent;\n-\t\t\t\t\tprint OUT \"Merge parent branch: $mparent\\n\" if $opt_v;\n-\t\t\t\t}\n-\t\t\t}\n+my($patchset,$date,$author_name,$author_email,$branch,$ancestor,$tag,$logmsg);\n+my(@old,@new,@skipped);\n+sub commit {\n+\tupdate_index(@old, @new);\n+\t@old = @new = ();\n+\tmy $tree = write_tree();\n+\tmy $parent = get_headref($last_branch, $git_dir);\n+\tprint \"Parent ID \" . ($parent ? $parent : \"(empty)\") . \"\\n\" if $opt_v;\n+\n+\tmy @commit_args;\n+\tpush @commit_args, (\"-p\", $parent) if $parent;\n+\n+\t# loose detection of merges\n+\t# based on the commit msg\n+\tforeach my $rx (@mergerx) {\n+\t\tnext unless $logmsg =~ $rx && $1;\n+\t\tmy $mparent = $1 eq 'HEAD' ? $opt_o : $1;\n+\t\tif(my $sha1 = get_headref($mparent, $git_dir)) {\n+\t\t\tpush @commit_args, '-p', $mparent;\n+\t\t\tprint \"Merge parent branch: $mparent\\n\" if $opt_v;\n \t\t}\n-\n-\t\texec(\"env\",\n-\t\t\t\"GIT_AUTHOR_NAME=$author_name\",\n-\t\t\t\"GIT_AUTHOR_EMAIL=$author_email\",\n-\t\t\t\"GIT_AUTHOR_DATE=\".strftime(\"+0000 %Y-%m-%d %H:%M:%S\",gmtime($date)),\n-\t\t\t\"GIT_COMMITTER_NAME=$author_name\",\n-\t\t\t\"GIT_COMMITTER_EMAIL=$author_email\",\n-\t\t\t\"GIT_COMMITTER_DATE=\".strftime(\"+0000 %Y-%m-%d %H:%M:%S\",gmtime($date)),\n-\t\t\t\"git-commit-tree\", $tree,@par);\n-\t\tdie \"Cannot exec git-commit-tree: $!\\n\";\n-\n-\t\tclose OUT;\n \t}\n-\t$pw->writer();\n-\t$pr->reader();\n+\n+\tmy $commit_date = strftime(\"+0000 %Y-%m-%d %H:%M:%S\",gmtime($date));\n+\tmy $pid = open2(my $commit_read, my $commit_write,\n+\t\t'env',\n+\t\t\"GIT_AUTHOR_NAME=$author_name\",\n+\t\t\"GIT_AUTHOR_EMAIL=$author_email\",\n+\t\t\"GIT_AUTHOR_DATE=$commit_date\",\n+\t\t\"GIT_COMMITTER_NAME=$author_name\",\n+\t\t\"GIT_COMMITTER_EMAIL=$author_email\",\n+\t\t\"GIT_COMMITTER_DATE=$commit_date\",\n+\t\t'git-commit-tree', $tree, @commit_args);\n \n \t# compatibility with git2cvs\n \tsubstr($logmsg,32767) = \"\" if length($logmsg) > 32767;\n@@ -656,16 +636,14 @@ sub commit {\n \t    @skipped = ();\n \t}\n \n-\tprint $pw \"$logmsg\\n\"\n+\tprint($commit_write \"$logmsg\\n\") && close($commit_write)\n \t\tor die \"Error writing to git-commit-tree: $!\\n\";\n-\t$pw->close();\n \n-\tprint \"Committed patch $patchset ($branch \".strftime(\"%Y-%m-%d %H:%M:%S\",gmtime($date)).\")\\n\" if $opt_v;\n-\tchomp(my $cid = <$pr>);\n-\tlength($cid) == 40\n-\t\tor die \"Cannot get commit id ($cid): $!\\n\";\n+\tprint \"Committed patch $patchset ($branch $commit_date)\\n\" if $opt_v;\n+\tchomp(my $cid = <$commit_read>);\n+\tis_sha1($cid) or die \"Cannot get commit id ($cid): $!\\n\";\n \tprint \"Commit ID $cid\\n\" if $opt_v;\n-\t$pr->close();\n+\tclose($commit_read);\n \n \twaitpid($pid,0);\n \tdie \"Error running git-commit-tree: $?\\n\" if $?;\n-- \n1.3.3.gcb64-dirty\n"},{"id":"20512","messageId":"20060523070102.GD6180@coredump.intra.peff.net","threadId":"4219","inReplyTo":"20060523065810.GB6180@coredump.intra.peff.net","subject":"[PATCH 1/2] cvsimport: use git-update-index --index-info","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T07:01:02Z","receivedAt":"2006-05-23T07:01:02Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This should reduce the number of git-update-index forks required per\ncommit. We now do adds/removes in one call, and we are no longer forced to\ndeal with argv limitations.\n\n---\n\nOops, apparently using a mail reader is too challenging for me. Here's a\nrepost with the headers correctly merged.\n\ncb6452bbfda9c52ad8dbeaac6e3440ae61099a05\n git-cvsimport.perl |   36 +++++++++++++-----------------------\n 1 files changed, 13 insertions(+), 23 deletions(-)\n\ncb6452bbfda9c52ad8dbeaac6e3440ae61099a05\ndiff --git a/git-cvsimport.perl b/git-cvsimport.perl\nindex d257e66..4efb0a5 100755\n--- a/git-cvsimport.perl\n+++ b/git-cvsimport.perl\n@@ -565,29 +565,19 @@ my($patchset,$date,$author_name,$author_\n my(@old,@new,@skipped);\n sub commit {\n \tmy $pid;\n-\twhile(@old) {\n-\t\tmy @o2;\n-\t\tif(@old > 55) {\n-\t\t\t@o2 = splice(@old,0,50);\n-\t\t} else {\n-\t\t\t@o2 = @old;\n-\t\t\t@old = ();\n-\t\t}\n-\t\tsystem(\"git-update-index\",\"--force-remove\",\"--\",@o2);\n-\t\tdie \"Cannot remove files: $?\\n\" if $?;\n-\t}\n-\twhile(@new) {\n-\t\tmy @n2;\n-\t\tif(@new > 12) {\n-\t\t\t@n2 = splice(@new,0,10);\n-\t\t} else {\n-\t\t\t@n2 = @new;\n-\t\t\t@new = ();\n-\t\t}\n-\t\tsystem(\"git-update-index\",\"--add\",\n-\t\t\t(map { ('--cacheinfo', @$_) } @n2));\n-\t\tdie \"Cannot add files: $?\\n\" if $?;\n-\t}\n+\n+      \topen(my $fh, '|-', qw(git-update-index --index-info))\n+\t\tor die \"unable to open git-update-index: $!\";\n+\tprint $fh \n+\t\t(map { \"0 0000000000000000000000000000000000000000\\t$_\\n\" }\n+\t\t\t@old),\n+\t\t(map { '100' . sprintf('%o', $_->[0]) . \" $_->[1]\\t$_->[2]\\n\" }\n+\t\t\t@new)\n+\t\tor die \"unable to write to git-update-index: $!\";\n+\tclose $fh\n+\t\tor die \"unable to write to git-update-index: $!\";\n+\t$? and die \"git-update-index reported error: $?\";\n+\t@old = @new = ();\n \n \t$pid = open(C,\"-|\");\n \tdie \"Cannot fork: $!\" unless defined $pid;\n-- \n1.3.3.gcb64-dirty\n"},{"id":"20513","messageId":"20060523071333.GA18249@coredump.intra.peff.net","threadId":"4219","inReplyTo":"7v4pzh6wtr.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T07:13:33Z","receivedAt":"2006-05-23T07:13:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"[cc'd to list to get reactions on open2]\n\nOn Tue, May 23, 2006 at 12:10:08AM -0700, Junio C Hamano wrote:\n\n> > +\treturn $s =~ /^[a-zA-Z0-9]{40}$/;\n> [0-9a-f] (We always do lowercase).\n\nEr, yes, that was a complete think-o on my part.\n\n> Hmm.  I personally do not have problems with open2, but folks on\n> some other platforms might.  I'll see how the list audience\n> sounds.\n\nFWIW, it was already being used in git-cvsimport.\n\n-Peff\n"},{"id":"20514","messageId":"37251.1135334664$1148369282@news.gmane.org","threadId":"4219","inReplyTo":"20060523070007.GC6180@coredump.intra.peff.net","subject":"[PATCH 1/2] cvsimport: use git-update-index --index-info","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T07:27:45Z","receivedAt":"2006-05-23T07:27:45Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This should reduce the number of git-update-index forks required per\ncommit. We now do adds/removes in one call, and we are no longer forced to\ndeal with argv limitations.\n\n---\n\nThis is a repost using -z/NUL instead of line feeds.\n\nd82d215430ae5e79210f73a31f5f8a053f36c27f\n git-cvsimport.perl |   36 +++++++++++++-----------------------\n 1 files changed, 13 insertions(+), 23 deletions(-)\n\nd82d215430ae5e79210f73a31f5f8a053f36c27f\ndiff --git a/git-cvsimport.perl b/git-cvsimport.perl\nindex d257e66..a65bea6 100755\n--- a/git-cvsimport.perl\n+++ b/git-cvsimport.perl\n@@ -565,29 +565,19 @@ my($patchset,$date,$author_name,$author_\n my(@old,@new,@skipped);\n sub commit {\n \tmy $pid;\n-\twhile(@old) {\n-\t\tmy @o2;\n-\t\tif(@old > 55) {\n-\t\t\t@o2 = splice(@old,0,50);\n-\t\t} else {\n-\t\t\t@o2 = @old;\n-\t\t\t@old = ();\n-\t\t}\n-\t\tsystem(\"git-update-index\",\"--force-remove\",\"--\",@o2);\n-\t\tdie \"Cannot remove files: $?\\n\" if $?;\n-\t}\n-\twhile(@new) {\n-\t\tmy @n2;\n-\t\tif(@new > 12) {\n-\t\t\t@n2 = splice(@new,0,10);\n-\t\t} else {\n-\t\t\t@n2 = @new;\n-\t\t\t@new = ();\n-\t\t}\n-\t\tsystem(\"git-update-index\",\"--add\",\n-\t\t\t(map { ('--cacheinfo', @$_) } @n2));\n-\t\tdie \"Cannot add files: $?\\n\" if $?;\n-\t}\n+\n+\topen(my $fh, '|-', qw(git-update-index -z --index-info))\n+\t\tor die \"unable to open git-update-index: $!\";\n+\tprint $fh \n+\t\t(map { \"0 0000000000000000000000000000000000000000\\t$_\\0\" }\n+\t\t\t@old),\n+\t\t(map { '100' . sprintf('%o', $_->[0]) . \" $_->[1]\\t$_->[2]\\0\" }\n+\t\t\t@new)\n+\t\tor die \"unable to write to git-update-index: $!\";\n+\tclose $fh\n+\t\tor die \"unable to write to git-update-index: $!\";\n+\t$? and die \"git-update-index reported error: $?\";\n+\t@old = @new = ();\n \n \t$pid = open(C,\"-|\");\n \tdie \"Cannot fork: $!\" unless defined $pid;\n-- \n1.3.3.g3408\n"},{"id":"20519","messageId":"46a038f90605230113x2f6b0e4bq5a2ea97308b495e0@mail.gmail.com","threadId":"4219","inReplyTo":"20060523070007.GC6180@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-23T08:13:10Z","receivedAt":"2006-05-23T08:13:10Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"Jeff,\n\ngood stuff -- aiming at exactly the things that had been nagging me.\nSome minor notes on top of what junio's mentioned...\n\n> +    die \"unable to open $f: $!\" unless $! == POSIX::ENOENT;\n> +    return undef;\n\nHeh. Is that the return of the living dead?\n\n> +sub update_index (\\@\\@) {\n> +       my $old = shift;\n> +       my $new = shift;\n\nWould it not make more sense to just pass them as plain parameters?\n\n> +       print \"Committed patch $patchset ($branch $commit_date)\\n\" if\n\nGiven that we have that -- should we remember it and avoid re-reading\nthe headref from disk? A %seenheads cache would save us 99.9% of the\nhassle.\n\nIn related news, I've dealt with file reads from the socket being\nmemorybound. Should merge ok.\n\ncheers,\n\n\nmartin\n"},{"id":"20521","messageId":"7vpsi55et5.fsf@assigned-by-dhcp.cox.net","threadId":"4219","inReplyTo":"46a038f90605230113x2f6b0e4bq5a2ea97308b495e0@mail.gmail.com","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-23T08:24:38Z","receivedAt":"2006-05-23T08:24:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> Jeff,\n>\n> good stuff -- aiming at exactly the things that had been nagging me.\n> Some minor notes on top of what junio's mentioned...\n>\n>> +    die \"unable to open $f: $!\" unless $! == POSIX::ENOENT;\n>> +    return undef;\n>\n> Heh. Is that the return of the living dead?\n\nNote the trailing \"unless\" there.\n\n>> +sub update_index (\\@\\@) {\n>> +       my $old = shift;\n>> +       my $new = shift;\n>\n> Would it not make more sense to just pass them as plain parameters?\n\nMeaning...?  Perl5 can pass only one flat array, so the above is\na standard way to pass two arrays.\n\n>> +       print \"Committed patch $patchset ($branch $commit_date)\\n\" if\n>\n> Given that we have that -- should we remember it and avoid re-reading\n> the headref from disk? A %seenheads cache would save us 99.9% of the\n> hassle.\n>\n> In related news, I've dealt with file reads from the socket being\n> memorybound. Should merge ok.\n\nMerged OK, and I think your last suggestion makes sense.  I'll\ngo to bed after pushing out Jeff's two patches and yours.\n"},{"id":"20556","messageId":"Pine.LNX.4.64.0605230948280.5623@g5.osdl.org","threadId":"4219","inReplyTo":"46a038f90605230113x2f6b0e4bq5a2ea97308b495e0@mail.gmail.com","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-23T16:50:01Z","receivedAt":"2006-05-23T16:50:01Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nHmm. Is it just me, or does the current \"git cvsimport\" have new problems:\n\n\t[torvalds@merom git]$ git cvsimport -d ~/CVS gentoo-x86\n\ncauses\n\n\tCommitting initial tree 34bd3dcd4bfd79bad35ce3fb08b2e21108195db8\n\tServer has gone away while fetching BUGS-TODO 1.1, retrying...\n\tRetry failed at /home/torvalds/bin/git-cvsimport line 366, <GEN2656> line 9.\n\nand that's it for the import.\n\nI don't see what would have caused it in the changes, but it definitely \nworked earlier..\n\n\t\tLinus\n"},{"id":"20557","messageId":"118833cc0605231047o2012deefh5e77b8496da1e673@mail.gmail.com","threadId":"4219","inReplyTo":"20060523070007.GC6180@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2006-05-23T17:47:01Z","receivedAt":"2006-05-23T17:47:01Z","isPatch":true,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"Why run \"env\" and not just muck with %ENV?\n\nM.\n\n\n> +       my $pid = open2(my $commit_read, my $commit_write,\n> +               'env',\n> +               \"GIT_AUTHOR_NAME=$author_name\",\n> +               \"GIT_AUTHOR_EMAIL=$author_email\",\n> +               \"GIT_AUTHOR_DATE=$commit_date\",\n> +               \"GIT_COMMITTER_NAME=$author_name\",\n> +               \"GIT_COMMITTER_EMAIL=$author_email\",\n> +               \"GIT_COMMITTER_DATE=$commit_date\",\n> +               'git-commit-tree', $tree, @commit_args);\n"},{"id":"20565","messageId":"Pine.LNX.4.64.0605231232360.5623@g5.osdl.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605230948280.5623@g5.osdl.org","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-23T19:36:37Z","receivedAt":"2006-05-23T19:36:37Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 23 May 2006, Linus Torvalds wrote:\n> \n> Hmm. Is it just me, or does the current \"git cvsimport\" have new problems:\n> \n> \t[torvalds@merom git]$ git cvsimport -d ~/CVS gentoo-x86\n> \n> causes\n> \n> \tCommitting initial tree 34bd3dcd4bfd79bad35ce3fb08b2e21108195db8\n> \tServer has gone away while fetching BUGS-TODO 1.1, retrying...\n> \tRetry failed at /home/torvalds/bin/git-cvsimport line 366, <GEN2656> line 9.\n> \n> and that's it for the import.\n> \n> I don't see what would have caused it in the changes, but it definitely \n> worked earlier..\n\nMartin, that problem seems to go away when I initialize $res to 0 in \n_fetchfile. \n\nI don't know perl, and maybe local variables are pre-initialized to empty. \n\nIt's entirely possible that the fact that it now seems to work for me is \npurely timing-related, since I also ended up using \"-P cvsps-output\" to \navoid having a huge cvsps binary in memory at the same time.\n\n\t\tLinus \"perl illiterate\" Torvalds\n"},{"id":"20573","messageId":"e4vqob$apj$1@sea.gmane.org","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605221055270.3697@g5.osdl.org","subject":"Re: irc usage..","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-05-23T20:19:33Z","receivedAt":"2006-05-23T20:19:33Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds wrote:\n \n> [...] people _should_ realize that removing objects is very very special. \n> Whether it's done by \"git prune-packed\" or \"git prune\", that's a very \n> dangerous operations. \"git prune\" a lot more so than \"git prune-packed\", \n> of course (in fact, you should _never_ run \"git prune\" on a repository \n> that is active - you _will_ corrupt it)-\n\nWould it be possible to make 'git prune' command repository corruption safe,\neven if some information might be lost (like 'git add')? Or do _corruption_\nmean some recoverable only information is lost? Not always one can use \"one\nrepository per developer\" workflow.\n\n\nOne of the solution would be to to use reader/writer lock (filesystem\nsemaphore), with each command modyfying repository performing locking, and\ngit-prune waiting on lock until noone is accessing repository. Of course\nthe problem is with OS and filesystems which does not support locking, and\nwith stale locks...\n\nSecond solution would be to [optionally] wait until no process is accessing\nrepository, copy repository in some safe place, [optionally] calculate\nchecksum, prune, [optionally] check if the repository was modified\nmeanwhile and either abort or repeat, and finally copy pruned repository\nback.\n\n-- \nJakub Narebski\nWarsaw, Poland\n"},{"id":"20574","messageId":"7vsln04hf0.fsf@assigned-by-dhcp.cox.net","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605231232360.5623@g5.osdl.org","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-23T20:25:55Z","receivedAt":"2006-05-23T20:25:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n>> \tCommitting initial tree 34bd3dcd4bfd79bad35ce3fb08b2e21108195db8\n>> \tServer has gone away while fetching BUGS-TODO 1.1, retrying...\n>...\n> Martin, that problem seems to go away when I initialize $res to 0 in \n> _fetchfile. \n>\n> I don't know perl, and maybe local variables are pre-initialized to empty. \n\nWhen a new file that is empty is created, sub _line would call\nsub _fetchfile with $cnt == 0, and it can return $res which\nis initialized to 'undef'.  That explains why sub file says\n$self->_line() returned an undef and I think what you did is the\nright fix.\n"},{"id":"20575","messageId":"46a038f90605231329w35d10cfdg1ac413ebf8d32e11@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605231232360.5623@g5.osdl.org","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-23T20:29:07Z","receivedAt":"2006-05-23T20:29:07Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/24/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> Martin, that problem seems to go away when I initialize $res to 0 in\n> _fetchfile.\n>\n> I don't know perl, and maybe local variables are pre-initialized to empty.\n>\n> It's entirely possible that the fact that it now seems to work for me is\n> purely timing-related, since I also ended up using \"-P cvsps-output\" to\n> avoid having a huge cvsps binary in memory at the same time.\n\nStrange! Cannot repro here with v5.8.8 (debian/etch 5.8.8-4) but\ninitialising it doesn't hurt, so let's do it:\n\ndiff --git a/git-cvsimport.perl b/git-cvsimport.perl\nindex ace7087..abbfd0b 100755\n--- a/git-cvsimport.perl\n+++ b/git-cvsimport.perl\n@@ -371,7 +371,7 @@ sub file {\n }\n sub _fetchfile {\n        my ($self, $fh, $cnt) = @_;\n-       my $res;\n+       my $res = 0;\n        my $bufsize = 1024 * 1024;\n        while($cnt) {\n            if ($bufsize > $cnt) {\n\ncheers,\n\n\nmartin\n"},{"id":"20576","messageId":"46a038f90605231332l7ba7a596k2916e6e8c7456e48@mail.gmail.com","threadId":"4219","inReplyTo":"7vpsi55et5.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-23T20:32:16Z","receivedAt":"2006-05-23T20:32:16Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/23/06, Junio C Hamano <junkio@cox.net> wrote:\n> \"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n>\n> > Jeff,\n> >\n> > good stuff -- aiming at exactly the things that had been nagging me.\n> > Some minor notes on top of what junio's mentioned...\n> >\n> >> +    die \"unable to open $f: $!\" unless $! == POSIX::ENOENT;\n> >> +    return undef;\n> >\n> > Heh. Is that the return of the living dead?\n>\n> Note the trailing \"unless\" there.\n\nOf course. I had actually missed the closing quotes, and thought the\nerror msg wanted to talk about POSIX. 'twas late in the day, seems\nlike most of my comments in this email were rather stoopid.\n\n> >> +sub update_index (\\@\\@) {\n> >> +       my $old = shift;\n> >> +       my $new = shift;\n> >\n> > Would it not make more sense to just pass them as plain parameters?\n>\n> Meaning...?  Perl5 can pass only one flat array, so the above is\n> a standard way to pass two arrays.\n\nMeaning I am stupid :(\n\n> >> +       print \"Committed patch $patchset ($branch $commit_date)\\n\" if\n> >\n> > Given that we have that -- should we remember it and avoid re-reading\n> > the headref from disk? A %seenheads cache would save us 99.9% of the\n> > hassle.\n> >\n> > In related news, I've dealt with file reads from the socket being\n> > memorybound. Should merge ok.\n>\n> Merged OK, and I think your last suggestion makes sense.  I'll\n> go to bed after pushing out Jeff's two patches and yours.\n\nI'll look into caching headrefs tonight if noone beats me to it.\n\n\n\n\nmartin\n"},{"id":"20580","messageId":"20060523205944.GA16164@coredump.intra.peff.net","threadId":"4219","inReplyTo":"118833cc0605231047o2012deefh5e77b8496da1e673@mail.gmail.com","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T20:59:44Z","receivedAt":"2006-05-23T20:59:44Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, May 23, 2006 at 01:47:01PM -0400, Morten Welinder wrote:\n\n> Why run \"env\" and not just muck with %ENV?\n> >+       my $pid = open2(my $commit_read, my $commit_write,\n> >+               'env',\n> >+               \"GIT_AUTHOR_NAME=$author_name\",\n> >+               \"GIT_AUTHOR_EMAIL=$author_email\",\n> >+               \"GIT_AUTHOR_DATE=$commit_date\",\n> >+               \"GIT_COMMITTER_NAME=$author_name\",\n> >+               \"GIT_COMMITTER_EMAIL=$author_email\",\n> >+               \"GIT_COMMITTER_DATE=$commit_date\",\n> >+               'git-commit-tree', $tree, @commit_args);\n\nOops, that's an obvious fork optimization that I should have caught.\nPatch is below. Note that this will now affect the environment of all\nsub-processes, but it shouldn't matter since we reset it right before\ncommit. However, if anyone is worried, we can stash the old %ENV in\nanother hash temporarily.\n\n-Peff\n\nPS What is the preferred format for throwing patches into replies like\nthis? Putting the patch at the end (as here) or throwing the reply\ncomments in the ignored section near the diffstat?\n\n---\ncvsimport: set up commit environment in perl instead of using env\n\n---\n\n44c4a9f67322302ca49146a7c143c07ea67da366\n git-cvsimport.perl |   13 ++++++-------\n 1 files changed, 6 insertions(+), 7 deletions(-)\n\n44c4a9f67322302ca49146a7c143c07ea67da366\ndiff --git a/git-cvsimport.perl b/git-cvsimport.perl\nindex 41ee9a6..83d7d3c 100755\n--- a/git-cvsimport.perl\n+++ b/git-cvsimport.perl\n@@ -618,14 +618,13 @@ sub commit {\n \t}\n \n \tmy $commit_date = strftime(\"+0000 %Y-%m-%d %H:%M:%S\",gmtime($date));\n+\t$ENV{GIT_AUTHOR_NAME} = $author_name;\n+\t$ENV{GIT_AUTHOR_EMAIL} = $author_email;\n+\t$ENV{GIT_AUTHOR_DATE} = $commit_date;\n+\t$ENV{GIT_COMMITTER_NAME} = $author_name;\n+\t$ENV{GIT_COMMITTER_EMAIL} = $author_email;\n+\t$ENV{GIT_COMMITTER_DATE} = $commit_date;\n \tmy $pid = open2(my $commit_read, my $commit_write,\n-\t\t'env',\n-\t\t\"GIT_AUTHOR_NAME=$author_name\",\n-\t\t\"GIT_AUTHOR_EMAIL=$author_email\",\n-\t\t\"GIT_AUTHOR_DATE=$commit_date\",\n-\t\t\"GIT_COMMITTER_NAME=$author_name\",\n-\t\t\"GIT_COMMITTER_EMAIL=$author_email\",\n-\t\t\"GIT_COMMITTER_DATE=$commit_date\",\n \t\t'git-commit-tree', $tree, @commit_args);\n \n \t# compatibility with git2cvs\n-- \n1.3.3.g40505-dirty\n\n\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"20582","messageId":"20060523211016.GB16164@coredump.intra.peff.net","threadId":"4219","inReplyTo":"46a038f90605231329w35d10cfdg1ac413ebf8d32e11@mail.gmail.com","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-23T21:10:16Z","receivedAt":"2006-05-23T21:10:16Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 24, 2006 at 08:29:07AM +1200, Martin Langhoff wrote:\n\n> Strange! Cannot repro here with v5.8.8 (debian/etch 5.8.8-4) but\n> initialising it doesn't hurt, so let's do it:\n\nI can reproduce with debian perl 5.8.8-4. The bug is only triggered by\n0-length files, so presumably your test repo doesn't have any.\n\n-Peff\n"},{"id":"20583","messageId":"46a038f90605231413g5f9f46a6p9cf15726526484f0@mail.gmail.com","threadId":"4219","inReplyTo":"20060523211016.GB16164@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-23T21:13:33Z","receivedAt":"2006-05-23T21:13:33Z","isPatch":true,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/24/06, Jeff King <peff@peff.net> wrote:\n> On Wed, May 24, 2006 at 08:29:07AM +1200, Martin Langhoff wrote:\n>\n> > Strange! Cannot repro here with v5.8.8 (debian/etch 5.8.8-4) but\n> > initialising it doesn't hurt, so let's do it:\n>\n> I can reproduce with debian perl 5.8.8-4. The bug is only triggered by\n> 0-length files, so presumably your test repo doesn't have any.\n\nGiven that we are all working off the gentoo repo here, it means that\nmy machine is slower than Linus' unreleased Intel box. And that I am\ntoo impatient...\n\nIn any case, the fix is correct as Junio points out.\n\ncheers,\n\n\nmartin\n"},{"id":"20591","messageId":"7vpsi41f82.fsf@assigned-by-dhcp.cox.net","threadId":"4219","inReplyTo":"20060523205944.GA16164@coredump.intra.peff.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-05-23T23:41:33Z","receivedAt":"2006-05-23T23:41:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Tue, May 23, 2006 at 01:47:01PM -0400, Morten Welinder wrote:\n>\n>> Why run \"env\" and not just muck with %ENV?\n>> >+       my $pid = open2(my $commit_read, my $commit_write,\n>> >+               'env',\n>> >+               \"GIT_AUTHOR_NAME=$author_name\",\n>> >+               \"GIT_AUTHOR_EMAIL=$author_email\",\n>> >+               \"GIT_AUTHOR_DATE=$commit_date\",\n>> >+               \"GIT_COMMITTER_NAME=$author_name\",\n>> >+               \"GIT_COMMITTER_EMAIL=$author_email\",\n>> >+               \"GIT_COMMITTER_DATE=$commit_date\",\n>> >+               'git-commit-tree', $tree, @commit_args);\n>\n> Oops, that's an obvious fork optimization that I should have caught.\n\nAre you two talking about running git-commit-tree via env is two\nfork-execs instead of just one?  Does that have a measurable\ndifference?\n\nNot that I have anything against the updated code, but I do not\nparticularly thing it is such a big issue.\n\n> PS What is the preferred format for throwing patches into replies like\n> this? Putting the patch at the end (as here) or throwing the reply\n> comments in the ignored section near the diffstat?\n\nYou could do it either way.  Although I personally find the\nformer easier to read (meshes well with \"do not top post\"\nmantra), it appears many other people finds the cover letter\nmaterial should come after the first '---' separator.\n\nIf you append the patch to your message, btw, you would need to\nrealize that the receiving end needs to edit your message to\nremove the top part before running \"git am\" to apply.\n"},{"id":"20618","messageId":"20060524095212.GA29510@coredump.intra.peff.net","threadId":"4219","inReplyTo":"7vpsi41f82.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH 2/2] cvsimport: cleanup commit function","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-05-24T09:52:12Z","receivedAt":"2006-05-24T09:52:12Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, May 23, 2006 at 04:41:33PM -0700, Junio C Hamano wrote:\n\n> Are you two talking about running git-commit-tree via env is two\n> fork-execs instead of just one?  Does that have a measurable\n> difference?\n\nYes, that's what I was talking about. No, probably not a huge\ndifference. I did some performance measurements of all of the recent\ncvsimport changes on a small-ish personal repo (I don't have the gentoo\nrepo). The results were not significant (<= 1% improvement for each\nchange).  I would expect some of the changes (index-info, fetchfile) to\nhave an impact on a repo with different characteristics (like the gentoo\none).\n\n-Peff\n"},{"id":"20922","messageId":"447B6D85.4050601@gentoo.org","threadId":"4219","inReplyTo":"447231C4.2030508@gentoo.org","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-29T21:54:13Z","receivedAt":"2006-05-29T21:54:13Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Donnie Berkholz wrote:\n> Linus Torvalds wrote:\n>> The latest stable CVS release is 1.11.21, I think: you seem to be running \n>> the \"development\" version (1.12.x).\n> \n> Backed down to the 1.11 series, things seem to be going fine so far.\n\nFinally hit an OOM sometime in the past day (yep, a week later) =\\. Not\nsure whether it was cvsimport or cvs. Anyone else had more luck?\n\nThanks,\nDonnie\n\n"},{"id":"20930","messageId":"46a038f90605291521q37f34209wd923608bdebb9084@mail.gmail.com","threadId":"4219","inReplyTo":"447B6D85.4050601@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-29T22:21:14Z","receivedAt":"2006-05-29T22:21:14Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> Donnie Berkholz wrote:\n> > Linus Torvalds wrote:\n> >> The latest stable CVS release is 1.11.21, I think: you seem to be running\n> >> the \"development\" version (1.12.x).\n> >\n> > Backed down to the 1.11 series, things seem to be going fine so far.\n>\n> Finally hit an OOM sometime in the past day (yep, a week later) =\\. Not\n> sure whether it was cvsimport or cvs. Anyone else had more luck?\n\nIt seemed like it had finished on the machine I was running it, and I\nassumed it was alright in yours too. Looking closer it only made it\ntill April 2004 -- but it may have been killed by a sysadmin, the\ncaptured log talks about 'signal 9', I have no idea what the OOM\nsends.\n\nIt had done 285070 of 343822 patchsets.\n\nHave you dropped the -a from the git-repack invocation? That should\nhelp. Try also Linus' patch for git-rev-list. The other thing hurting\nus is that the commits are _huge_. I wonder how you guys were managing\nthis with CVS. Now _this_ explains why cvsimport grows humongous.\n\nI'll try to rework the commit loop so that we don't need to hold all\nthe filenames in memory. It seems to be choking with the commits after\nApril 2004. But that will have to wait till tonight.\n\ncheers,\n\n\n\nmartin\n"},{"id":"20933","messageId":"447B7669.8050805@gentoo.org","threadId":"4219","inReplyTo":"46a038f90605291521q37f34209wd923608bdebb9084@mail.gmail.com","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-29T22:32:09Z","receivedAt":"2006-05-29T22:32:09Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n>> Finally hit an OOM sometime in the past day (yep, a week later) =\\. Not\n>> sure whether it was cvsimport or cvs. Anyone else had more luck?\n> \n> It seemed like it had finished on the machine I was running it, and I\n> assumed it was alright in yours too. Looking closer it only made it\n> till April 2004 -- but it may have been killed by a sysadmin, the\n> captured log talks about 'signal 9', I have no idea what the OOM\n> sends.\n\nLooking closer, I see that the memory suckers do appear to be git, from\ndmesg:\n\nOut of Memory: Kill process 17230 (git-repack) score 97207 and children.\nOut of memory: Killed process 17231 (git-rev-list).\n\nJust ends like this:\n\nTree ID 2cc632e5e1d3a430a2cc891bf33c4a12f19a4d0e\nParent ID ad92d7073a52458e0581633bbd8ccbbec838d9e6\nCommitted patch 249100 (origin 2005-08-20 05:05:58)\nCommit ID 28941f00d714f57ab49f1fd725d1c3ce8a5d0b93\nFetching sys-kernel/ck-sources/ChangeLog   v 1.113\nUpdate sys-kernel/ck-sources/ChangeLog: 25425 bytes\nFetching sys-kernel/ck-sources/Manifest   v 1.164\nUpdate sys-kernel/ck-sources/Manifest: 252 bytes\nDelete sys-kernel/ck-sources/ck-sources-2.6.12_p5-r1.ebuild\nFetching sys-kernel/ck-sources/ck-sources-2.6.12_p6.ebuild   v 1.1\nNew sys-kernel/ck-sources/ck-sources-2.6.12_p6.ebuild: 1438 bytes\nDelete sys-kernel/ck-sources/files/digest-ck-sources-2.6.12_p5-r1\nFetching sys-kernel/ck-sources/files/digest-ck-sources-2.6.12_p6   v 1.1\nNew sys-kernel/ck-sources/files/digest-ck-sources-2.6.12_p6: 279 bytes\nCan't fork at /usr/bin/git-cvsimport line 592, <CVS> line 3810053.\ncat: write error: Broken pipe\n\n> It had done 285070 of 343822 patchsets.\n> \n> Have you dropped the -a from the git-repack invocation? That should\n> help. Try also Linus' patch for git-rev-list. The other thing hurting\n> us is that the commits are _huge_. I wonder how you guys were managing\n> this with CVS. Now _this_ explains why cvsimport grows humongous.\n\nI wasn't running with a version that did repacks; I just suspended the\ncvsimport a couple of times and ran a repack manually.\n\n> I'll try to rework the commit loop so that we don't need to hold all\n> the filenames in memory. It seems to be choking with the commits after\n> April 2004. But that will have to wait till tonight.\n\nThanks,\nDonnie\n\n"},{"id":"20941","messageId":"46a038f90605291719r292269bct61bf2817a9791e3d@mail.gmail.com","threadId":"4219","inReplyTo":"447B7669.8050805@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-30T00:19:20Z","receivedAt":"2006-05-30T00:19:20Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> Looking closer, I see that the memory suckers do appear to be git, from\n> dmesg:\n>\n> Out of Memory: Kill process 17230 (git-repack) score 97207 and children.\n> Out of memory: Killed process 17231 (git-rev-list).\n\nThat would mean that you do have Linus' patch then. Grep cvsimport for\nrepack and remove the -a -- and consider using his recent patch to\nrev-list.\n\nMy dmesg talks about an earlier cvs segfault. Nasty tree you have here\n-- it's breaking all sorts of things... and teaching us a thing or two\nabout the import process.\n\n> Committed patch 249100 (origin 2005-08-20 05:05:58)\n\nHmmm? How can you be at patch 249100 and still be a good year ahead of\nme? Have you told cvsps to cut off old history?\n\nAnother thing I found is that this import uses a lot of $TMPDIR, so if\nyour TMPDIR is small, you'll hit all sorts of problems.\n\ncheers,\n\n\n\nmartin\n"},{"id":"20942","messageId":"Pine.LNX.4.64.0605291742520.5623@g5.osdl.org","threadId":"4219","inReplyTo":"447B7669.8050805@gentoo.org","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-30T00:43:57Z","receivedAt":"2006-05-30T00:43:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 29 May 2006, Donnie Berkholz wrote:\n> \n> Looking closer, I see that the memory suckers do appear to be git, from\n> dmesg:\n> \n> Out of Memory: Kill process 17230 (git-repack) score 97207 and children.\n> Out of memory: Killed process 17231 (git-rev-list).\n\nSounds like you had the \"git repack -a -d\" thing in your cvsimport.\n\nThe current git rev-list should use only about a third of the memory of \nthe one you used, so hopefully you could just update your git version, and \nthen continue with the \"git cvsimport\" without having to start all over.\n\n\t\tLinus\n"},{"id":"20956","messageId":"447BD8C1.6090402@gentoo.org","threadId":"4219","inReplyTo":"46a038f90605291719r292269bct61bf2817a9791e3d@mail.gmail.com","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-30T05:31:45Z","receivedAt":"2006-05-30T05:31:45Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n>> Looking closer, I see that the memory suckers do appear to be git, from\n>> dmesg:\n>>\n>> Out of Memory: Kill process 17230 (git-repack) score 97207 and children.\n>> Out of memory: Killed process 17231 (git-rev-list).\n> \n> That would mean that you do have Linus' patch then. Grep cvsimport for\n> repack and remove the -a -- and consider using his recent patch to\n> rev-list.\n\nYou certainly would think so, and I did as well, but available evidence\nindicates otherwise. I'm not sure how the repack got in there.\n\ndonnie@supernova ~ $ type git-cvsimport\ngit-cvsimport is /usr/bin/git-cvsimport\ndonnie@supernova ~ $ grep repack /usr/bin/git-cvsimport\ndonnie@supernova ~ $\n\nAll I can think of is that I somehow OOM'd when I manually ran a repack\nand didn't notice it. But that should've at least made me unable to\nresume the cvsimport process, which happily kept chugging along later on.\n\n> My dmesg talks about an earlier cvs segfault. Nasty tree you have here\n> -- it's breaking all sorts of things... and teaching us a thing or two\n> about the import process.\n> \n>> Committed patch 249100 (origin 2005-08-20 05:05:58)\n> \n> Hmmm? How can you be at patch 249100 and still be a good year ahead of\n> me? Have you told cvsps to cut off old history?\n\nNope. I ran the exact cvsps flags you posted earlier to create it.\n\nThanks,\nDonnie\n\n"},{"id":"20958","messageId":"46a038f90605292301h667291cdo7260683f8e746933@mail.gmail.com","threadId":"4219","inReplyTo":"447BD8C1.6090402@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-30T06:01:58Z","receivedAt":"2006-05-30T06:01:58Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> All I can think of is that I somehow OOM'd when I manually ran a repack\n> and didn't notice it. But that should've at least made me unable to\n> resume the cvsimport process, which happily kept chugging along later on.\n\nSounds likely -- and cvsimport restarts gracefully, though you might want to do\n\n   git checkout HEAD\n\nto get a usable checkout if the very first import failed. However, the\ndefault head is master, and what you want to look at is origin or\nwhatever you passed as your -o parameter. I use cvshead normally, so I\ndo\n\n   git log cvshead\n\n> > My dmesg talks about an earlier cvs segfault. Nasty tree you have here\n> > -- it's breaking all sorts of things... and teaching us a thing or two\n> > about the import process.\n> >\n> >> Committed patch 249100 (origin 2005-08-20 05:05:58)\n> >\n> > Hmmm? How can you be at patch 249100 and still be a good year ahead of\n> > me? Have you told cvsps to cut off old history?\n>\n> Nope. I ran the exact cvsps flags you posted earlier to create it.\n\nOh, that was an earlier PEBKAK at my end: I did git log HEAD instead\nof git log cvshead. My import is now at  293145 (cvshead +0000\n2005-12-25 12:24:42) which looks promising.\n\ncheers,\n\n\nmartin\n"},{"id":"20988","messageId":"46a038f90605301531g4f8b37c7qab9a717833c64ebc@mail.gmail.com","threadId":"4219","inReplyTo":"447B7669.8050805@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-30T22:31:35Z","receivedAt":"2006-05-30T22:31:35Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> Martin Langhoff wrote:\n> > On 5/30/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> >> Finally hit an OOM sometime in the past day (yep, a week later) =\\. Not\n> >> sure whether it was cvsimport or cvs. Anyone else had more luck?\n\nWith the latest cvsimport in Junio's repo, a lot of RAM and a bit of patience...\n\n  gitview\n  http://git.catalyst.net.nz/gitweb?p=gentoo.git;a=summary\n\n  fetchable\n  http://git.catalyst.net.nz/git/gentoo.git#cvshead\n\nStill pushing it, will be there in a minute or so. The packed repo\nweights about 660MB. Not too bad given the size of the project and the\nnumber of commits.\n\n\nmartin\n"},{"id":"20990","messageId":"Pine.LNX.4.64.0605301604130.24646@g5.osdl.org","threadId":"4219","inReplyTo":"46a038f90605301531g4f8b37c7qab9a717833c64ebc@mail.gmail.com","subject":"Re: irc usage..","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-05-30T23:07:47Z","receivedAt":"2006-05-30T23:07:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 31 May 2006, Martin Langhoff wrote:\n> \n>  gitview\n>  http://git.catalyst.net.nz/gitweb?p=gentoo.git;a=summary\n\nHeh. I think you should enable caching in your apache config. \n\nAnd maybe we should make that part of the gitweb docs. Without a caching \nweb-server, gitweb is pretty slow, but it caches _beautifully_.\n\nThat gentoo repo has a lot of \"duplicate\" commits that cvsps will mark as \ntwo separate commits because there's one commit for the files, and one \ncommit for whatever the \"Manifest\" file is. I wonder if those commits \nshould generally be merged or something. \n\nThat said, things like that are most easily fixed as a git->git update \n(along with adding name translation), which can avoid re-writing the \ntrees.\n\n\t\t\tLinus\n"},{"id":"20993","messageId":"46a038f90605301804u3beabf4ct97c8a0ea6ef7b995@mail.gmail.com","threadId":"4219","inReplyTo":"Pine.LNX.4.64.0605301604130.24646@g5.osdl.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-31T01:04:46Z","receivedAt":"2006-05-31T01:04:46Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/31/06, Linus Torvalds <torvalds@osdl.org> wrote:\n> On Wed, 31 May 2006, Martin Langhoff wrote:\n> >\n> >  gitview\n> >  http://git.catalyst.net.nz/gitweb?p=gentoo.git;a=summary\n>\n> Heh. I think you should enable caching in your apache config.\n\nI know I should -- but I'm hoping to find the time to rework gitweb a\nbit to actually work fast instead. It bothers me that it is so slow on\na basically idle machine, and where I can perform the corresponding\ngit operations in the commandline in a blink.\n\nAnd caching is great for really busy sites (aka kernel.org) but\ngit.catalyst.net.nz only serves a handful of small repos for a small\ngroup of people, and is 99% idle. Should blaze through this stuff.\n\n> That gentoo repo has a lot of \"duplicate\" commits that cvsps will mark as\n> two separate commits because there's one commit for the files, and one\n> commit for whatever the \"Manifest\" file is. I wonder if those commits\n> should generally be merged or something.\n>\n> That said, things like that are most easily fixed as a git->git update\n> (along with adding name translation), which can avoid re-writing the\n> trees.\n\nYep, large projects often have good reasons to run custom imports,\nmerging certain commits, rewriting log messages (like the X.org guys\nwere doing). It can be done at the cvsimport stage or later -- I think\nPasky has a rewritehistory tool hidden somewhere in Cogito, but I\nhaven't used it.\n\ncheers,\n\n\nmartin\n"},{"id":"20996","messageId":"447D043D.1020609@gentoo.org","threadId":"4219","inReplyTo":"46a038f90605301804u3beabf4ct97c8a0ea6ef7b995@mail.gmail.com","subject":"Re: irc usage..","fromName":"Donnie Berkholz","fromEmail":"spyderous@gentoo.org","sentAt":"2006-05-31T02:49:33Z","receivedAt":"2006-05-31T02:49:33Z","isPatch":false,"sender":{"key":"spyderous@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 5/31/06, Linus Torvalds <torvalds@osdl.org> wrote:\n>> That gentoo repo has a lot of \"duplicate\" commits that cvsps will mark as\n>> two separate commits because there's one commit for the files, and one\n>> commit for whatever the \"Manifest\" file is. I wonder if those commits\n>> should generally be merged or something.\n>>\n>> That said, things like that are most easily fixed as a git->git update\n>> (along with adding name translation), which can avoid re-writing the\n>> trees.\n> \n> Yep, large projects often have good reasons to run custom imports,\n> merging certain commits, rewriting log messages (like the X.org guys\n> were doing). It can be done at the cvsimport stage or later -- I think\n> Pasky has a rewritehistory tool hidden somewhere in Cogito, but I\n> haven't used it.\n\nWe've got a guy who got a Summer of Code project to work on CVS\nmigration, so this could be something along his lines.\n\nThanks,\nDonnie\n\n"},{"id":"21000","messageId":"46a038f90605302305g7a969a62r277af1724b912069@mail.gmail.com","threadId":"4219","inReplyTo":"447D043D.1020609@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-31T06:05:53Z","receivedAt":"2006-05-31T06:05:53Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/31/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n> We've got a guy who got a Summer of Code project to work on CVS\n> migration, so this could be something along his lines.\n\nHe'll want a fast box to wrangle with this repo ;-)\n\n\nmartin\n"},{"id":"21004","messageId":"447DA028.3040606@gentoo.org","threadId":"4219","inReplyTo":"46a038f90605302305g7a969a62r277af1724b912069@mail.gmail.com","subject":"Re: irc usage..","fromName":"Alec Warner","fromEmail":"antarus@gentoo.org","sentAt":"2006-05-31T13:54:48Z","receivedAt":"2006-05-31T13:54:48Z","isPatch":false,"sender":{"key":"antarus@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 5/31/06, Donnie Berkholz <spyderous@gentoo.org> wrote:\n>> We've got a guy who got a Summer of Code project to work on CVS\n>> migration, so this could be something along his lines.\n> \n> He'll want a fast box to wrangle with this repo ;-)\n> \n> \n> martin\n\nI have a dual opteron with 4gb of ram \"on loan\" from work :)\n\nIt still dies though, using git cvsimport or parsecvs.\n\nI talked to Keith Packard about adding support to parsecvs for recording \nthe actual changed changesets, but I haven't yet started on implementing \nthat since he isn't using cvsps in parsecvs.\n\nI also haven't had a chance to look at the git-cvsimport sources yet, \nwas hoping to get to that later this week.\n"},{"id":"21034","messageId":"46a038f90605311503o1526c664qe61b0f3f40929b92@mail.gmail.com","threadId":"4219","inReplyTo":"447DA028.3040606@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-05-31T22:03:44Z","receivedAt":"2006-05-31T22:03:44Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/1/06, Alec Warner <antarus@gentoo.org> wrote:\n> I have a dual opteron with 4gb of ram \"on loan\" from work :)\n>\n> It still dies though, using git cvsimport or parsecvs.\n\nThe machine I am running this is more constrained than that, and it\ndoesn't die. It just takes maybe 30hs. Make sure it's not a bad cvs\nbinary you got there (latest from gentoo seems to leak memory).\n\nAnd if it's still dying... give us some more details ;-)\n\ncheers,\n\n\nmartin\n"},{"id":"21042","messageId":"447E4611.7000309@gentoo.org","threadId":"4219","inReplyTo":"46a038f90605311503o1526c664qe61b0f3f40929b92@mail.gmail.com","subject":"Re: irc usage..","fromName":"Alec Warner","fromEmail":"antarus@gentoo.org","sentAt":"2006-06-01T01:42:41Z","receivedAt":"2006-06-01T01:42:41Z","isPatch":false,"sender":{"key":"antarus@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 6/1/06, Alec Warner <antarus@gentoo.org> wrote:\n> \n>> I have a dual opteron with 4gb of ram \"on loan\" from work :)\n>>\n>> It still dies though, using git cvsimport or parsecvs.\n> \n> \n> The machine I am running this is more constrained than that, and it\n> doesn't die. It just takes maybe 30hs. Make sure it's not a bad cvs\n> binary you got there (latest from gentoo seems to leak memory).\n> \n> And if it's still dying... give us some more details ;-)\n> \n> cheers,\n> \n> \n> martin\n\nAfter reading the whole thread on this, I've using a git checkout of \ngit, cvsps-2.1 and cvs-1.11.12, running overnight in verbose mode with \nscreen.  Hopefully will have a repo in the morning ;)\n"},{"id":"21047","messageId":"46a038f90606010047r676840d2nd91ad2361abbe1c8@mail.gmail.com","threadId":"4219","inReplyTo":"447E4611.7000309@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-01T07:47:03Z","receivedAt":"2006-06-01T07:47:03Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/1/06, Alec Warner <antarus@gentoo.org> wrote:\n> After reading the whole thread on this, I've using a git checkout of\n> git, cvsps-2.1 and cvs-1.11.12, running overnight in verbose mode with\n> screen.  Hopefully will have a repo in the morning ;)\n\nGood stuff. I am rerunning it to prove (and bench) a complete an\nuninterrupted import. So far it's done 4hs 30m, footprint grown to\n207MB, 49750 commits. So I think it will be done in approx 30hs on\nthis single-cpu opteron.\n\nMost commits are small, but there is a handful that are downright\nmassive -- and we hold all the file list in memory, which I think\nexplains (most of) the memory growth. I've looked into avoiding\nholding the whole filelist in memory, but it involves rewriting the\ncvsps output parsing loop, which is better left for a rainy day, with\na test case that doesn't take 30hs to resolve.\n\ncheers,\n\n\n\nmartin\n"},{"id":"21241","messageId":"44837BDB.2090601@gentoo.org","threadId":"4219","inReplyTo":"46a038f90606010047r676840d2nd91ad2361abbe1c8@mail.gmail.com","subject":"Re: irc usage..","fromName":"Alec Warner","fromEmail":"antarus@gentoo.org","sentAt":"2006-06-05T00:33:31Z","receivedAt":"2006-06-05T00:33:31Z","isPatch":false,"sender":{"key":"antarus@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 6/1/06, Alec Warner <antarus@gentoo.org> wrote:\n> \n>> After reading the whole thread on this, I've using a git checkout of\n>> git, cvsps-2.1 and cvs-1.11.12, running overnight in verbose mode with\n>> screen.  Hopefully will have a repo in the morning ;)\n> \n> \n> Good stuff. I am rerunning it to prove (and bench) a complete an\n> uninterrupted import. So far it's done 4hs 30m, footprint grown to\n> 207MB, 49750 commits. So I think it will be done in approx 30hs on\n> this single-cpu opteron.\n> \n> Most commits are small, but there is a handful that are downright\n> massive -- and we hold all the file list in memory, which I think\n> explains (most of) the memory growth. I've looked into avoiding\n> holding the whole filelist in memory, but it involves rewriting the\n> cvsps output parsing loop, which is better left for a rainy day, with\n> a test case that doesn't take 30hs to resolve.\n\nOk the box this was running on had issues, so I switched to using \npearl.amd64.dev.gentoo.org, a dual core amd64 X2 4600+ with 4 gigs of \nram and plenty of disk.  The \"problem\" now is just converstion time...30 \nhours and I'm into 2004-09-17...but it's been in 2004 all day, seems \nlike most of the commits are in the last three years.  Are there \narchitectural issues with doing this in parallel?\n\nSince the repository commits are all in cvs, it should be possible to do \nthe work in parallel, since you know what all the commits touch.  The \nconcern would be ordering of nodes in the tree; you'd end up building a \nbunch of subtrees and patching them together?\n\n-Alec Warner\n"},{"id":"21245","messageId":"46a038f90606041906k66d85152v6e402c65151d7ab8@mail.gmail.com","threadId":"4219","inReplyTo":"44837BDB.2090601@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-05T02:06:59Z","receivedAt":"2006-06-05T02:06:59Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/5/06, Alec Warner <antarus@gentoo.org> wrote:\n> Ok the box this was running on had issues, so I switched to using\n> pearl.amd64.dev.gentoo.org, a dual core amd64 X2 4600+ with 4 gigs of\n> ram and plenty of disk.  The \"problem\" now is just converstion time...30\n> hours and I'm into 2004-09-17...but it's been in 2004 all day, seems\n> like most of the commits are in the last three years.  Are there\n> architectural issues with doing this in parallel?\n\nI don't think you can do this in parallel. What I would do is remove\nthe -a from the git-repack invocation. It does hurt import times quite\na bit -- just do a git-repack -a -d when it's done.\n\nAnd... having said that, there is still a memory leak somehow,\nsomewhere. It's been evading me for 2 weeks now, so I feel an idiot\nnow. Not too bad in general, but it shows clearly in the gentoo and\nmozilla imports.\n\n> Since the repository commits are all in cvs, it should be possible to do\n> the work in parallel, since you know what all the commits touch.  The\n> concern would be ordering of nodes in the tree; you'd end up building a\n> bunch of subtrees and patching them together?\n\nWell... parsecvs does a bit of this but in sequential fashion... it\nimports all the files first, and then runs through the history\nbuilding the tree+commits in order, committing them. It saves a lot of\ntime in the file imports by parsing the RCS file directly. The\ndownside is that it must keep a filename+version=>sha1 mapping --\nwhich I think is why parsecvs won't fit in memory until it's changed\nto store it on disk somehow ;-)\n\nYou are forced to do it in a sequence because cvsps only tells you\nabout the files added/removed/changed in a commit -- you need the\nancestor to have a view of what the whole tree looked like. The only\nroom for parallelism I see is to fork off new processes to work on\nbranches in parallel.\n\n\n\nmartin\n"},{"id":"21249","messageId":"448398BC.5090402@gentoo.org","threadId":"4219","inReplyTo":"46a038f90606041906k66d85152v6e402c65151d7ab8@mail.gmail.com","subject":"Re: irc usage..","fromName":"Alec Warner","fromEmail":"antarus@gentoo.org","sentAt":"2006-06-05T02:36:44Z","receivedAt":"2006-06-05T02:36:44Z","isPatch":false,"sender":{"key":"antarus@gentoo.org","avatar":null},"body":"Martin Langhoff wrote:\n> On 6/5/06, Alec Warner <antarus@gentoo.org> wrote:\n> \n>> Ok the box this was running on had issues, so I switched to using\n>> pearl.amd64.dev.gentoo.org, a dual core amd64 X2 4600+ with 4 gigs of\n>> ram and plenty of disk.  The \"problem\" now is just converstion time...30\n>> hours and I'm into 2004-09-17...but it's been in 2004 all day, seems\n>> like most of the commits are in the last three years.  Are there\n>> architectural issues with doing this in parallel?\n> \n> \n> I don't think you can do this in parallel. What I would do is remove\n> the -a from the git-repack invocation. It does hurt import times quite\n> a bit -- just do a git-repack -a -d when it's done.\n\nOnly repack at the end then? disk space isn't an issue here so I'll give \nthat a shot.\n\n> \n> And... having said that, there is still a memory leak somehow,\n> somewhere. It's been evading me for 2 weeks now, so I feel an idiot\n> now. Not too bad in general, but it shows clearly in the gentoo and\n> mozilla imports.\n\n30565 antarus   17   0  470m 456m 1640 S   14 11.6 234:23.38\ngit-cvsimport\n30566 antarus   16   0 6753m 147m  752 S    7  3.7 120:27.06 cvs\n\nI'm on cvs-1.11.12 and the git version of git\n\n> You are forced to do it in a sequence because cvsps only tells you\n> about the files added/removed/changed in a commit -- you need the\n> ancestor to have a view of what the whole tree looked like. The only\n> room for parallelism I see is to fork off new processes to work on\n> branches in parallel.\n\nNot helpful in the Gentoo case, since we only have one branch; minus an \naccident when a dev branched gentoo-x86 a while back ;)\n\nI'll keep chugging on this one; it won't be the final import as I \nhaven't used the complete Authors file, so I will try the repacking \noptimization next time I do an import.\n\n-Alec Warner\n"},{"id":"21251","messageId":"46a038f90606042049y3dfb1bbdwc91132ddd9eeaa39@mail.gmail.com","threadId":"4219","inReplyTo":"448398BC.5090402@gentoo.org","subject":"Re: irc usage..","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-05T03:49:38Z","receivedAt":"2006-06-05T03:49:38Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/5/06, Alec Warner <antarus@gentoo.org> wrote:\n> > I don't think you can do this in parallel. What I would do is remove\n> > the -a from the git-repack invocation. It does hurt import times quite\n> > a bit -- just do a git-repack -a -d when it's done.\n>\n> Only repack at the end then? disk space isn't an issue here so I'll give\n> that a shot.\n\nNot exactly -- by removing the -a from the git-repack invocation what\nyou get is cheap \"partial\" packing rather than a full repack. This is\nsomewhat inefficient disk-wise, perhaps by 10% or so. But full repacks\nget more and more expensive as the repo grows.\n\nSo you don't need to run git-repack -a -d at the end, but it will be a\ngood measure to see how compact the packing gets.\n\n> > And... having said that, there is still a memory leak somehow,\n> > somewhere. It's been evading me for 2 weeks now, so I feel an idiot\n> > now. Not too bad in general, but it shows clearly in the gentoo and\n> > mozilla imports.\n>\n> 30565 antarus   17   0  470m 456m 1640 S   14 11.6 234:23.38\n> git-cvsimport\n> 30566 antarus   16   0 6753m 147m  752 S    7  3.7 120:27.06 cvs\n>\n> I'm on cvs-1.11.12 and the git version of git\n\nYep, I see roughly the same. It grows slowly and I don't know why :(\n\n> I'll keep chugging on this one; it won't be the final import as I\n> haven't used the complete Authors file, so I will try the repacking\n> optimization next time I do an import.\n\nCool. If it dies for any reason, just do\n\n  git-update-ref refs/heads/master refs/heads/origin\n  git-update-ref HEAD origin\n  git-checkout\n\nYou only need to do this the first time -- after that, the core heads\nare set. Rerun the script and it will pick up where it left. If it\ndies again, just do git-checkout to see the latest files.\n\n(Above, replace origin with your -o option if you are using it. I\nnormally use -o cvshead.)\n\n\n\nmartin\n"},{"id":"21260","messageId":"BAYC1-PASMTP029677186C792C538C1921AE940@CEZ.ICE","threadId":"4219","inReplyTo":"448398BC.5090402@gentoo.org","subject":"Re: irc usage..","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2006-06-05T16:07:43Z","receivedAt":"2006-06-05T16:07:43Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Sun, 04 Jun 2006 22:36:44 -0400\nAlec Warner <antarus@gentoo.org> wrote:\n\n> I'll keep chugging on this one; it won't be the final import as I \n> haven't used the complete Authors file, so I will try the repacking \n> optimization next time I do an import.\n\nHi Alec,\n\nYou may want to go back and do another import for other reasons, but if\nthe only reason is to fix up the author information it would be _much_\nfaster to simply rewrite the git commit history.  Cogito has something\ncalled \"cg-admin-rewritehist\" which should do what you need and there\nare other scripts floating around specificially for rewriting just the\nauthor information.\n\nHTH,\nSean\n"}]}