{"thread":{"id":"11153","subject":"Re: Git and GCC","startedAt":"2007-12-06T02:28:15Z","lastAt":"2009-03-18T18:02:39Z","messageCount":89,"participants":["David Miller","Daniel Berlin","Harvey Harrison","Linus Torvalds","Jon Smirl","Jeff King","David Brown","Johannes Schindelin","Ismail Dönmez","Theodore Tso","Nicolas Pitre","Pierre Habouzit","David Kastrup","NightStrike","Jon Loeliger","Junio C Hamano","Randy Dunlap","Jakub Narebski","Giovanni Bajo","Luke Lu","Gabriel Paubert","Geert Bosch","Derek Fawcus","Shawn O. Pearce","Teemu Likonen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"62084","messageId":"20071205.182815.249974508.davem@davemloft.net","threadId":"11153","inReplyTo":"4aca3dc20712051108s216d3331t8061ef45b9aa324a@mail.gmail.com","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-06T02:28:15Z","receivedAt":"2007-12-06T02:28:15Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: \"Daniel Berlin\" <dberlin@dberlin.org>\nDate: Wed, 5 Dec 2007 14:08:41 -0500\n\n> So I tried a full history conversion using git-svn of the gcc\n> repository (IE every trunk revision from 1-HEAD as of yesterday)\n> The git-svn import was done using repacks every 1000 revisions.\n> After it finished, I used git-gc --aggressive --prune.  Two hours\n> later, it finished.\n> The final size after this is 1.5 gig for all of the history of gcc for\n> just trunk.\n> \n> dberlin@home:/compilerstuff/gitgcc/gccrepo/.git/objects/pack$ ls -trl\n> total 1568899\n> -r--r--r-- 1 dberlin dberlin 1585972834 2007-12-05 14:01\n> pack-cd328fcf0bd673d8f2f72c42fbe67da64cbcd218.pack\n> -r--r--r-- 1 dberlin dberlin   19008488 2007-12-05 14:01\n> pack-cd328fcf0bd673d8f2f72c42fbe67da64cbcd218.idx\n> \n> This is 3x bigger than hg *and* hg doesn't require me to waste my life\n> repacking every so often.\n> The hg operations run roughly as fast as the git ones\n> \n> I'm sure there are magic options, magic command lines, etc, i could\n> use to make it smaller.\n> \n> I'm sure if i spent the next few weeks fucking around with git, it may\n> even be usable!\n> \n> But given that git is harder to use, requires manual repacking to get\n> any kind of sane space usage, and is 3x bigger anyway, i don't see any\n> advantage to continuing to experiment with git and gcc.\n\nI would really appreciate it if you would share experiences\nlike this with the GIT community, who have been now CC:'d.\n\nThat's the only way this situation is going to improve.\n\nWhen you don't CC: the people who can fix the problem, I can only\nspeculate that perhaps at least subconsciously you don't care if\nthe situation improves or not.\n\nThe OpenSolaris folks behaved similarly, and that really ticked me\noff.\n"},{"id":"62085","messageId":"4aca3dc20712051841o71ab773ft6dd0714ebc355dd5@mail.gmail.com","threadId":"11153","inReplyTo":"20071205.182815.249974508.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-06T02:41:19Z","receivedAt":"2007-12-06T02:41:19Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> Date: Wed, 5 Dec 2007 14:08:41 -0500\n>\n> > So I tried a full history conversion using git-svn of the gcc\n> > repository (IE every trunk revision from 1-HEAD as of yesterday)\n> > The git-svn import was done using repacks every 1000 revisions.\n> > After it finished, I used git-gc --aggressive --prune.  Two hours\n> > later, it finished.\n> > The final size after this is 1.5 gig for all of the history of gcc for\n> > just trunk.\n> >\n> > dberlin@home:/compilerstuff/gitgcc/gccrepo/.git/objects/pack$ ls -trl\n> > total 1568899\n> > -r--r--r-- 1 dberlin dberlin 1585972834 2007-12-05 14:01\n> > pack-cd328fcf0bd673d8f2f72c42fbe67da64cbcd218.pack\n> > -r--r--r-- 1 dberlin dberlin   19008488 2007-12-05 14:01\n> > pack-cd328fcf0bd673d8f2f72c42fbe67da64cbcd218.idx\n> >\n> > This is 3x bigger than hg *and* hg doesn't require me to waste my life\n> > repacking every so often.\n> > The hg operations run roughly as fast as the git ones\n> >\n> > I'm sure there are magic options, magic command lines, etc, i could\n> > use to make it smaller.\n> >\n> > I'm sure if i spent the next few weeks fucking around with git, it may\n> > even be usable!\n> >\n> > But given that git is harder to use, requires manual repacking to get\n> > any kind of sane space usage, and is 3x bigger anyway, i don't see any\n> > advantage to continuing to experiment with git and gcc.\n>\n> I would really appreciate it if you would share experiences\n> like this with the GIT community, who have been now CC:'d.\n>\n> That's the only way this situation is going to improve.\n>\n> When you don't CC: the people who can fix the problem, I can only\n> speculate that perhaps at least subconsciously you don't care if\n> the situation improves or not.\n>\nI didn't cc the git community for three reasons\n\n1. It's not the nicest message in the world, and thus, more likely to\nget bad responses than constructive ones.\n\n2. Based on the level of usability, I simply assume it is too young\nfor regular developers to use.  At least, I hope this is the case.\n\n3. People i know have had bad experiences talking usability issues\nwith the git community in the past.  I am not likely to fare any\nbetter, so I would rather have someone who is involved with both our\ncommunity and theirs, raise these issues, rather than a complete\nnewcomer.\n\nBut hey, whatever floats your boat :)\n\nIt is true I gave up quickly, but this is mainly because i don't like\nto fight with my tools.\nI am quite fine with a distributed workflow, I now use 8 or so gcc\nbranches in mercurial (auto synced from svn) and merge a lot between\nthem. I wanted to see if git would sanely let me manage the commits\nback to svn.  After fighting with it, i gave up and just wrote a\npython extension to hg that lets me commit non-svn changesets back to\nsvn directly from hg.\n\n--Dan\n"},{"id":"62086","messageId":"20071205.185203.262588544.davem@davemloft.net","threadId":"11153","inReplyTo":"4aca3dc20712051841o71ab773ft6dd0714ebc355dd5@mail.gmail.com","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-06T02:52:03Z","receivedAt":"2007-12-06T02:52:03Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: \"Daniel Berlin\" <dberlin@dberlin.org>\nDate: Wed, 5 Dec 2007 21:41:19 -0500\n\n> It is true I gave up quickly, but this is mainly because i don't like\n> to fight with my tools.\n> I am quite fine with a distributed workflow, I now use 8 or so gcc\n> branches in mercurial (auto synced from svn) and merge a lot between\n> them. I wanted to see if git would sanely let me manage the commits\n> back to svn.  After fighting with it, i gave up and just wrote a\n> python extension to hg that lets me commit non-svn changesets back to\n> svn directly from hg.\n\nI find it ironic that you were even willing to write tools to\nfacilitate your hg based gcc workflow.  That really shows what your\nthinking is on this matter, in that you're willing to put effort\ntowards making hg work better for you but you're not willing to expend\nthat level of effort to see if git can do so as well.\n\nThis is what really eats me from the inside about your dissatisfaction\nwith git.  Your analysis seems to be a self-fullfilling prophecy, and\nthat's totally unfair to both hg and git.\n"},{"id":"62088","messageId":"4aca3dc20712051947t5fbbb383ua1727c652eb25d7e@mail.gmail.com","threadId":"11153","inReplyTo":"20071205.185203.262588544.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-06T03:47:01Z","receivedAt":"2007-12-06T03:47:01Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> Date: Wed, 5 Dec 2007 21:41:19 -0500\n>\n> > It is true I gave up quickly, but this is mainly because i don't like\n> > to fight with my tools.\n> > I am quite fine with a distributed workflow, I now use 8 or so gcc\n> > branches in mercurial (auto synced from svn) and merge a lot between\n> > them. I wanted to see if git would sanely let me manage the commits\n> > back to svn.  After fighting with it, i gave up and just wrote a\n> > python extension to hg that lets me commit non-svn changesets back to\n> > svn directly from hg.\n>\n> I find it ironic that you were even willing to write tools to\n> facilitate your hg based gcc workflow.\nWhy?\n\n> That really shows what your\n> thinking is on this matter, in that you're willing to put effort\n> towards making hg work better for you but you're not willing to expend\n> that level of effort to see if git can do so as well.\nSee, now you claim to know my thinking.\nI went back to hg because the GIT's space usage wasn't even in the\nballpark, i couldn't get git-svn rebase to update the revs after the\ninitial import (even though i had properly used a rewriteRoot).\n\nThe size is clearly not just svn data, it's in the git pack itself.\n\nI spent a long time working on SVN to reduce it's space usage (repo\nside and cleaning up the client side and giving a path to svn devs to\nreduce it further), as well as ui issues, and I really don't feel like\nhaving to do the same for GIT.\n\nI'm tired of having to spend a large amount of effort to get my tools\nto work.  If the community wants to find and fix the problem, i've\nalready said repeatedly i'll happily give over my repo, data,\nwhatever.  You are correct i am not going to spend even more effort\nwhen i can be productive with something else much quicker.  The devil\ni know (committing to svn) is better than the devil i don't (diving\ninto git source code and finding/fixing what is causing this space\nblowup).\nThe python extension took me a few hours (< 4).\nIn git, i spent these hours waiting for git-gc to finish.\n\n> This is what really eats me from the inside about your dissatisfaction\n> with git.  Your analysis seems to be a self-fullfilling prophecy, and\n> that's totally unfair to both hg and git.\nOh?\nYou seem to be taking this awfully personally.\nI came into this completely open minded. Really, I did (i'm sure\nyou'll claim otherwise).\nGIT people told me it would work great and i'd have a really small git\nrepo and be able to commit back to svn.\nI tried it.\nIt didn't work out.\nIt doesn't seem to be usable for whatever reason.\nI'm happy to give details, data, whatever.\n\nI made the engineering decision that my effort would be better spent\ndoing something I knew i could do quickly (make hg commit back to svn\nfor my purposes) then trying to improve larger issues in GIT (UI and\nspace usage).  That took me a few hours, and I was happy again.\n\nI would have been incredibly happy to have git just have come up with\na 400 meg gcc repository, and to be happily committing away from\ngit-svn to gcc's repository  ...\nBut it didn't happen.\nSo far, you have yet to actually do anything but incorrectly tell me\nwhat I am thinking.\n\nI'll probably try again in 6 months, and maybe it will be better.\n"},{"id":"62089","messageId":"20071205.202047.58135920.davem@davemloft.net","threadId":"11153","inReplyTo":"4aca3dc20712051947t5fbbb383ua1727c652eb25d7e@mail.gmail.com","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-06T04:20:47Z","receivedAt":"2007-12-06T04:20:47Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: \"Daniel Berlin\" <dberlin@dberlin.org>\nDate: Wed, 5 Dec 2007 22:47:01 -0500\n\n> The size is clearly not just svn data, it's in the git pack itself.\n\nAnd other users have shown much smaller metadata from a GIT import,\nand yes those are including all of the repository history and branches\nnot just the trunk.\n"},{"id":"62090","messageId":"1196915112.10408.66.camel@brick","threadId":"11153","inReplyTo":"4aca3dc20712051947t5fbbb383ua1727c652eb25d7e@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-06T04:25:07Z","receivedAt":"2007-12-06T04:25:07Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"I fought with this a few months ago when I did my own clone of gcc svn.\nMy bad for only discussing this on #git at the time.  Should have put\nthis to the list as well.\n\nIf anyone recalls my report was something along the lines of\ngit gc --aggressive explodes pack size.\n\ngit repack -a -d --depth=100 --window=100 produced a ~550MB packfile\nimmediately afterwards a git gc --aggressive produces a 1.5G packfile.\n\nThis was for all branches/tags, not just trunk like Daniel's repo.\n\nThe best theory I had at the time was that the gc doesn't find as good\ndeltas or doesn't allow the same delta chain depth and so generates a \nnew object in the pack, rather the reusing a good delta it already has\nin the well-packed pack.\n\nCheers,\n\nHarvey\n"},{"id":"62091","messageId":"1196915319.10408.71.camel@brick","threadId":"11153","inReplyTo":"20071205.202047.58135920.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-06T04:28:39Z","receivedAt":"2007-12-06T04:28:39Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"\nOn Wed, 2007-12-05 at 20:20 -0800, David Miller wrote:\n> From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> Date: Wed, 5 Dec 2007 22:47:01 -0500\n> \n> > The size is clearly not just svn data, it's in the git pack itself.\n> \n> And other users have shown much smaller metadata from a GIT import,\n> and yes those are including all of the repository history and branches\n> not just the trunk.\n\nDavid, I think it is actually a bug in git gc with the --aggressive\noption...mind you, even if he solves that the format git svn uses\nfor its bi-directional metadata is so space-inefficient Daniel will\nbe crying for other reasons immediately afterwards...4MB for every\nbranch and tag in gcc svn (more than a few thousand).\n\nYou only need it around for any branches you are planning on committing\nto but it is all created during the default git svn import.\n\nFYI\n\nHarvey\n"},{"id":"62092","messageId":"4aca3dc20712052032n521c344cla07a5df1f2c26cb8@mail.gmail.com","threadId":"11153","inReplyTo":"20071205.202047.58135920.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-06T04:32:52Z","receivedAt":"2007-12-06T04:32:52Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> Date: Wed, 5 Dec 2007 22:47:01 -0500\n>\n> > The size is clearly not just svn data, it's in the git pack itself.\n>\n> And other users have shown much smaller metadata from a GIT import,\n> and yes those are including all of the repository history and branches\n> not just the trunk.\nI followed the instructions in the tutorials.\nI followed the instructions given to by people who created these.\nI came up with a 1.5 gig pack file.\nYou want to help, or you want to argue with me.\nRight now it sounds like you are trying to blame me or make it look\nlike i did something wrong.\n\nYou are of course, welcome to try it yourself.\nI can give you the absolute exactly commands I gave, and with git\n1.5.3.7, it will give you a 1.5 gig pack file.\n"},{"id":"62095","messageId":"20071205.204848.227521641.davem@davemloft.net","threadId":"11153","inReplyTo":"4aca3dc20712052032n521c344cla07a5df1f2c26cb8@mail.gmail.com","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-06T04:48:48Z","receivedAt":"2007-12-06T04:48:48Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: \"Daniel Berlin\" <dberlin@dberlin.org>\nDate: Wed, 5 Dec 2007 23:32:52 -0500\n\n> On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> > From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> > Date: Wed, 5 Dec 2007 22:47:01 -0500\n> >\n> > > The size is clearly not just svn data, it's in the git pack itself.\n> >\n> > And other users have shown much smaller metadata from a GIT import,\n> > and yes those are including all of the repository history and branches\n> > not just the trunk.\n> I followed the instructions in the tutorials.\n> I followed the instructions given to by people who created these.\n> I came up with a 1.5 gig pack file.\n> You want to help, or you want to argue with me.\n\nSeveral people replied in this thread showing what options can lead to\nsmaller pack files.\n\nThey also listed what the GIT limitations are that would effect the\nkind of work you are doing, which seemed to mostly deal with the high\nspace cost of branching and tags when converting to/from SVN repos.\n"},{"id":"62098","messageId":"alpine.LFD.0.9999.0712052033570.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"1196915112.10408.66.camel@brick","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T04:54:28Z","receivedAt":"2007-12-06T04:54:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Dec 2007, Harvey Harrison wrote:\n> \n> If anyone recalls my report was something along the lines of\n> git gc --aggressive explodes pack size.\n\nYes, --aggressive is generally a bad idea. I think we should remove it or \nat least fix it. It doesn't do what the name implies, because it actually \nthrows away potentially good packing, and re-does it all from a clean \nslate.\n\nThat said, it's totally pointless for a person who isn't a git proponent \nto do an initial import, and in that sense I agree with Daniel: he \nshouldn't waste his time with tools that he doesn't know or care about, \nsince there are people who *can* do a better job, and who know what they \nare doing, and understand and like the tool.\n\nWhile you can do a half-assed job with just mindlessly running \"git \nsvnimport\" (which is deprecated these days) or \"git svn clone\" (better), \nthe fact is, to do a *good* import does likely mean spending some effort \non it. Trying to make the user names / emails to be better with a mailmap, \nfor example. \n\n[ By default, for example, \"git svn clone/fetch\" seems to create those \n  horrible fake email addresses that contain the ID of the SVN repo in \n  each commit - I'm not talking about the \"git-svn-id\", I'm talking about \n  the \"user@hex-string-goes-here\" thing for the author. Maybe people don't \n  really care, but isn't that ugly as hell? I'd think it's worth it doing \n  a really nice import, spending some effort on it.\n\n  But maybe those things come from the older CVS->SVN import, I don't \n  really know. I've done a few SVN imports, but I've done them just for \n  stuff where I didn't want to touch SVN, but just wanted to track some \n  project like libgpod. For things like *that*, a totally mindless \"git \n  svn\" thing is fine ]\n\nOf course, that does require there to be git people in the gcc crowd who \nare motivated enough to do the proper import and then make sure it's \nup-to-date and hosted somewhere. If those people don't exist, I'm not sure \nthere's much idea to it.\n\nThe point being, you cannot ask a non-git person to do a major git import \nfor an actual switch-over. Yes, it *can* be as simple as just doing a\n\n\tgit svn clone --stdlayout svn://svn://gcc.gnu.org/svn/gcc gcc\n\nbut the fact remains, you want to spend more effort and expertise on it if \nyou actually want the result to be used as a basis for future work (as \nopposed to just tracking somebody elses SVN tree).\n\nThat includes:\n\n - do the historic import with good packing (and no, \"--aggressive\" \n   is not it, never mind the misleading name and man-page)\n\n - probably mailmap entries, certainly spending some time validating the \n   results.\n\n - hosting it\n\nand perhaps most importantly\n\n - helping people who are *not* git users get up to speed.\n\nbecause doing a good job at it is like asking a CVS newbie to set up a \nbranch in CVS. I'm sure you can do it from man-pages, but I'm also sure \nyou sure as hell won't like the end result.\n\n\t\tLinus\n"},{"id":"62099","messageId":"1196917487.10408.82.camel@brick","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712052033570.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-06T05:04:47Z","receivedAt":"2007-12-06T05:04:47Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Wed, 2007-12-05 at 20:54 -0800, Linus Torvalds wrote:\n> \n> On Wed, 5 Dec 2007, Harvey Harrison wrote:\n> > \n> > If anyone recalls my report was something along the lines of\n> > git gc --aggressive explodes pack size.\n\n> [ By default, for example, \"git svn clone/fetch\" seems to create those \n>   horrible fake email addresses that contain the ID of the SVN repo in \n>   each commit - I'm not talking about the \"git-svn-id\", I'm talking about \n>   the \"user@hex-string-goes-here\" thing for the author. Maybe people don't \n>   really care, but isn't that ugly as hell? I'd think it's worth it doing \n>   a really nice import, spending some effort on it.\n> \n>   But maybe those things come from the older CVS->SVN import, I don't \n>   really know. I've done a few SVN imports, but I've done them just for \n>   stuff where I didn't want to touch SVN, but just wanted to track some \n>   project like libgpod. For things like *that*, a totally mindless \"git \n>   svn\" thing is fine ]\n> \n\ngit svn does accept a mailmap at import time with the same format as the\ncvs importer I think.  But for someone that just wants a repo to check\nout this was easiest.  I'd be willing to spend the time to do a nicer\njob if there was any interest from the gcc side, but I'm not that\ninvested (other than owing them for an often-used tool).\n\nHarvey\n"},{"id":"62100","messageId":"4aca3dc20712052111o730f6fb6h7a329ee811a70f28@mail.gmail.com","threadId":"11153","inReplyTo":"20071205.204848.227521641.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-06T05:11:05Z","receivedAt":"2007-12-06T05:11:05Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> Date: Wed, 5 Dec 2007 23:32:52 -0500\n>\n> > On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> > > From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> > > Date: Wed, 5 Dec 2007 22:47:01 -0500\n> > >\n> > > > The size is clearly not just svn data, it's in the git pack itself.\n> > >\n> > > And other users have shown much smaller metadata from a GIT import,\n> > > and yes those are including all of the repository history and branches\n> > > not just the trunk.\n> > I followed the instructions in the tutorials.\n> > I followed the instructions given to by people who created these.\n> > I came up with a 1.5 gig pack file.\n> > You want to help, or you want to argue with me.\n>\n> Several people replied in this thread showing what options can lead to\n> smaller pack files.\n\nActually, one person did, but that's okay, let's assume it was several.\nI am currently trying Harvey's options.\n\nI asked about using the pre-existing repos so i didn't have to do\nthis, but they were all\n1. Done using read-only imports or\n2. Don't contain full history\n(IE the one that contains full history that is often posted here was\ndone as a read only import and thus doesn't have the metadata).\n\n> They also listed what the GIT limitations are that would effect the\n> kind of work you are doing, which seemed to mostly deal with the high\n> space cost of branching and tags when converting to/from SVN repos.\n\nActually, it turns out that git-gc --aggressive does this dumb thing\nto pack files sometimes regardless of whether you converted from an\nSVN repo or not.\n"},{"id":"62101","messageId":"1196918132.10408.85.camel@brick","threadId":"11153","inReplyTo":"4aca3dc20712052111o730f6fb6h7a329ee811a70f28@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-06T05:15:32Z","receivedAt":"2007-12-06T05:15:32Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Thu, 2007-12-06 at 00:11 -0500, Daniel Berlin wrote:\n> On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> > From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> > Date: Wed, 5 Dec 2007 23:32:52 -0500\n> >\n> > > On 12/5/07, David Miller <davem@davemloft.net> wrote:\n> > > > From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> > > > Date: Wed, 5 Dec 2007 22:47:01 -0500\n> > > >\n> > > > > The size is clearly not just svn data, it's in the git pack itself.\n> > > >\n> > > > And other users have shown much smaller metadata from a GIT import,\n> > > > and yes those are including all of the repository history and branches\n> > > > not just the trunk.\n> > > I followed the instructions in the tutorials.\n> > > I followed the instructions given to by people who created these.\n> > > I came up with a 1.5 gig pack file.\n> > > You want to help, or you want to argue with me.\n> >\n> > Several people replied in this thread showing what options can lead to\n> > smaller pack files.\n> \n> Actually, one person did, but that's okay, let's assume it was several.\n> I am currently trying Harvey's options.\n> \n> I asked about using the pre-existing repos so i didn't have to do\n> this, but they were all\n> 1. Done using read-only imports or\n> 2. Don't contain full history\n> (IE the one that contains full history that is often posted here was\n> done as a read only import and thus doesn't have the metadata).\n\nWhile you won't get the git svn metadata if you clone the infradead\nrepo, it can be recreated on the fly by git svn if you want to start\ncommiting directly to gcc svn.\n\nHarvey\n"},{"id":"62103","messageId":"4aca3dc20712052117j3ef5cf99y848d4962ae8ddf33@mail.gmail.com","threadId":"11153","inReplyTo":"1196918132.10408.85.camel@brick","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-06T05:17:11Z","receivedAt":"2007-12-06T05:17:11Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"> While you won't get the git svn metadata if you clone the infradead\n> repo, it can be recreated on the fly by git svn if you want to start\n> commiting directly to gcc svn.\n>\nI will give this a try :)\n"},{"id":"62110","messageId":"alpine.LFD.0.9999.0712052132450.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"4aca3dc20712052111o730f6fb6h7a329ee811a70f28@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T06:09:12Z","receivedAt":"2007-12-06T06:09:12Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Daniel Berlin wrote:\n> \n> Actually, it turns out that git-gc --aggressive does this dumb thing\n> to pack files sometimes regardless of whether you converted from an\n> SVN repo or not.\n\nAbsolutely. git --aggressive is mostly dumb. It's really only useful for \nthe case of \"I know I have a *really* bad pack, and I want to throw away \nall the bad packing decisions I have done\".\n\nTo explain this, it's worth explaining (you are probably aware of it, but \nlet me go through the basics anyway) how git delta-chains work, and how \nthey are so different from most other systems.\n\nIn other SCM's, a delta-chain is generally fixed. It might be \"forwards\" \nor \"backwards\", and it might evolve a bit as you work with the repository, \nbut generally it's a chain of changes to a single file represented as some \nkind of single SCM entity. In CVS, it's obviously the *,v file, and a lot \nof other systems do rather similar things.\n\nGit also does delta-chains, but it does them a lot more \"loosely\". There \nis no fixed entity. Delta's are generated against any random other version \nthat git deems to be a good delta candidate (with various fairly \nsuccessful heursitics), and there are absolutely no hard grouping rules.\n\nThis is generally a very good thing. It's good for various conceptual \nreasons (ie git internally never really even needs to care about the whole \nrevision chain - it doesn't really think in terms of deltas at all), but \nit's also great because getting rid of the inflexible delta rules means \nthat git doesn't have any problems at all with merging two files together, \nfor example - there simply are no arbitrary *,v \"revision files\" that have \nsome hidden meaning.\n\nIt also means that the choice of deltas is a much more open-ended \nquestion. If you limit the delta chain to just one file, you really don't \nhave a lot of choices on what to do about deltas, but in git, it really \ncan be a totally different issue.\n\nAnd this is where the really badly named \"--aggressive\" comes in. While \ngit generally tries to re-use delta information (because it's a good idea, \nand it doesn't waste CPU time re-finding all the good deltas we found \nearlier), sometimes you want to say \"let's start all over, with a blank \nslate, and ignore all the previous delta information, and try to generate \na new set of deltas\".\n\nSo \"--aggressive\" is not really about being aggressive, but about wasting \nCPU time re-doing a decision we already did earlier!\n\n*Sometimes* that is a good thing. Some import tools in particular could \ngenerate really horribly bad deltas. Anything that uses \"git fast-import\", \nfor example, likely doesn't have much of a great delta layout, so it might \nbe worth saying \"I want to start from a clean slate\".\n\nBut almost always, in other cases, it's actually a really bad thing to do. \nIt's going to waste CPU time, and especially if you had actually done a \ngood job at deltaing earlier, the end result isn't going to re-use all \nthose *good* deltas you already found, so you'll actually end up with a \nmuch worse end result too!\n\nI'll send a patch to Junio to just remove the \"git gc --aggressive\" \ndocumentation. It can be useful, but it generally is useful only when you \nreally understand at a very deep level what it's doing, and that \ndocumentation doesn't help you do that.\n\nGenerally, doing incremental \"git gc\" is the right approach, and better \nthan doing \"git gc --aggressive\". It's going to re-use old deltas, and \nwhen those old deltas can't be found (the reason for doing incremental GC \nin the first place!) it's going to create new ones.\n\nOn the other hand, it's definitely true that an \"initial import of a long \nand involved history\" is a point where it can be worth spending a lot of \ntime finding the *really*good* deltas. Then, every user ever after (as \nlong as they don't use \"git gc --aggressive\" to undo it!) will get the \nadvantage of that one-time event. So especially for big projects with a \nlong history, it's probably worth doing some extra work, telling the delta \nfinding code to go wild.\n\nSo the equivalent of \"git gc --aggressive\" - but done *properly* - is to \ndo (overnight) something like\n\n\tgit repack -a -d --depth=250 --window=250\n\nwhere that depth thing is just about how deep the delta chains can be \n(make them longer for old history - it's worth the space overhead), and \nthe window thing is about how big an object window we want each delta \ncandidate to scan.\n\nAnd here, you might well want to add the \"-f\" flag (which is the \"drop all \nold deltas\", since you now are actually trying to make sure that this one \nactually finds good candidates.\n\nAnd then it's going to take forever and a day (ie a \"do it overnight\" \nthing). But the end result is that everybody downstream from that \nrepository will get much better packs, without having to spend any effort \non it themselves.\n\n\t\t\tLinus\n"},{"id":"62122","messageId":"9e4733910712052247x116cabb4q48ebafffb93f7e03@mail.gmail.com","threadId":"11153","inReplyTo":"4aca3dc20712052117j3ef5cf99y848d4962ae8ddf33@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-06T06:47:54Z","receivedAt":"2007-12-06T06:47:54Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Daniel Berlin <dberlin@dberlin.org> wrote:\n> > While you won't get the git svn metadata if you clone the infradead\n> > repo, it can be recreated on the fly by git svn if you want to start\n> > commiting directly to gcc svn.\n> >\n> I will give this a try :)\n\nBack when I was working on the Mozilla repository we were able to\nconvert the full 4GB CVS repository complete with all history into a\n450MB pack file. That work is where the git-fastimport tool came from.\nBut it took a month of messing with the import tools to achieve this\nand Mozilla still chose another VCS (mainly because of poor Windows\nsupport in git).\n\nLike Linus says, this type of command will yield the smallest pack file:\n git repack -a -d --depth=250 --window=250\n\nI do agree that importing multi-gigabyte repositories is not a daily\noccurrence nor a turn-key operation. There are significant issues when\ntranslating from one VCS to another. The lack of global branch\ntracking in CVS causes extreme problems on import. Hand editing of CVS\nfiles also caused endless trouble.\n\nThe key to converting repositories of this size is RAM. 4GB minimum,\nmore would be better. git-repack is not multi-threaded. There were a\nfew attempts at making it multi-threaded but none were too successful.\nIf I remember right, with loads of RAM, a repack on a 450MB repository\nwas taking about five hours on a 2.8Ghz Core2. But this is something\nyou only have to do once for the import. Later repacks will reuse the\noriginal deltas.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62124","messageId":"20071206071503.GA19504@coredump.intra.peff.net","threadId":"11153","inReplyTo":"9e4733910712052247x116cabb4q48ebafffb93f7e03@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-06T07:15:03Z","receivedAt":"2007-12-06T07:15:03Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 06, 2007 at 01:47:54AM -0500, Jon Smirl wrote:\n\n> The key to converting repositories of this size is RAM. 4GB minimum,\n> more would be better. git-repack is not multi-threaded. There were a\n> few attempts at making it multi-threaded but none were too successful.\n> If I remember right, with loads of RAM, a repack on a 450MB repository\n> was taking about five hours on a 2.8Ghz Core2. But this is something\n> you only have to do once for the import. Later repacks will reuse the\n> original deltas.\n\nActually, Nicolas put quite a bit of work into multi-threading the\nrepack process; the results have been in master for some time, and will\nbe in the soon-to-be-released v1.5.4.\n\nThe downside is that the threading partitions the object space, so the\nresulting size is not necessarily as small (but I don't know that\nanybody has done testing on large repos to find out how large the\ndifference is).\n\n-Peff\n"},{"id":"62127","messageId":"1196927361.13109.1.camel@brick","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712052132450.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-06T07:49:21Z","receivedAt":"2007-12-06T07:49:21Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"\n> \tgit repack -a -d --depth=250 --window=250\n> \n\nSince I have the whole gcc repo locally I'll give this a shot overnight\njust to see what can be done at the extreme end or things.\n\nHarvey\n"},{"id":"62130","messageId":"20071206081139.GA14370@old.davidb.org","threadId":"11153","inReplyTo":"1196927361.13109.1.camel@brick","subject":"Re: Git and GCC","fromName":"David Brown","fromEmail":"git@davidb.org","sentAt":"2007-12-06T08:11:39Z","receivedAt":"2007-12-06T08:11:39Z","isPatch":false,"sender":{"key":"git@davidb.org","avatar":"https://gravatar.com/avatar/94c86a2938470a74c2eac5e2b69afc0871f79a660295c02219597aba8cb101c1?d=mp&s=160"},"body":"On Wed, Dec 05, 2007 at 11:49:21PM -0800, Harvey Harrison wrote:\n>\n>> \tgit repack -a -d --depth=250 --window=250\n>> \n>\n>Since I have the whole gcc repo locally I'll give this a shot overnight\n>just to see what can be done at the extreme end or things.\n\nWhen I tried this on a very large repo, at least one with some large files\nin it, git quickly exceeded my physical memory and started thrashing the\nmachine.  I had good results with\n\n  git config pack.deltaCacheSize 512m\n  git config pack.windowMemory 512m\n\nof course adjusting based on your physical memory.  I think changing the\nwindowMemory will affect the resulting compression, so changing these\nratios might get better compression out of the result.\n\nIf you're really patient, though, you could leave the unbounded window,\nhope you have enough swap, and just let it run.\n\nDave\n"},{"id":"62134","messageId":"Pine.LNX.4.64.0712061150590.27959@racer.site","threadId":"11153","inReplyTo":"20071205.185203.262588544.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-12-06T11:57:06Z","receivedAt":"2007-12-06T11:57:06Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 5 Dec 2007, David Miller wrote:\n\n> From: \"Daniel Berlin\" <dberlin@dberlin.org>\n> Date: Wed, 5 Dec 2007 21:41:19 -0500\n> \n> > It is true I gave up quickly, but this is mainly because i don't like \n> > to fight with my tools.\n> >\n> > I am quite fine with a distributed workflow, I now use 8 or so gcc \n> > branches in mercurial (auto synced from svn) and merge a lot between \n> > them. I wanted to see if git would sanely let me manage the commits \n> > back to svn.  After fighting with it, i gave up and just wrote a \n> > python extension to hg that lets me commit non-svn changesets back to \n> > svn directly from hg.\n> \n> I find it ironic that you were even willing to write tools to facilitate \n> your hg based gcc workflow.  That really shows what your thinking is on \n> this matter, in that you're willing to put effort towards making hg work \n> better for you but you're not willing to expend that level of effort to \n> see if git can do so as well.\n\nWhile this is true...\n\n> This is what really eats me from the inside about your dissatisfaction \n> with git.  Your analysis seems to be a self-fullfilling prophecy, and \n> that's totally unfair to both hg and git.\n\n... I actually appreciate people complaining -- in the meantime.  It shows \nright away what group you belong to in the \"Those who can do, do, those \nwho can't, complain.\".\n\nYou can see that very easily on the git list, or on the #git channel on \nirc.freenode.net.  There is enough data for a study which yearns to be \nwritten, that shows how quickly we resolve issues with people that are \nsincerely interested in a solution.\n\n(Of course, on the other hand, there are also quite a few cases which show \nhow frustrating (for both sides) and unfruitful discussions started by a \ncomplaint are.)\n\nSo I fully expect an issue like Daniel's to be resolved in a matter of \nminutes on the git list, if the OP gives us a chance.  If we are not even \nCc'ed, you are completely right, she or he probably does not want the \nissue to be resolved.\n\nCiao,\nDscho\n"},{"id":"62136","messageId":"Pine.LNX.4.64.0712061201580.27959@racer.site","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712052132450.13796@woody.linux-foundation.org","subject":"[PATCH] gc --aggressive: make it really aggressive","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-12-06T12:03:38Z","receivedAt":"2007-12-06T12:03:38Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"\nThe default was not to change the window or depth at all.  As suggested\nby Jon Smirl, Linus Torvalds and others, default to\n\n\t--window=250 --depth=250\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n\n\tOn Wed, 5 Dec 2007, Linus Torvalds wrote:\n\n\t> On Thu, 6 Dec 2007, Daniel Berlin wrote:\n\t> > \n\t> > Actually, it turns out that git-gc --aggressive does this dumb \n\t> > thing to pack files sometimes regardless of whether you \n\t> > converted from an SVN repo or not.\n\t> \n\t> Absolutely. git --aggressive is mostly dumb. It's really only \n\t> useful for the case of \"I know I have a *really* bad pack, and I \n\t> want to throw away all the bad packing decisions I have done\".\n\t>\n\t> [...]\n\t> \n\t> So the equivalent of \"git gc --aggressive\" - but done *properly* \n\t> - is to do (overnight) something like\n\t> \n\t> \tgit repack -a -d --depth=250 --window=250\n\n\tHow about this, then?\n\t\n builtin-gc.c |    3 ++-\n 1 files changed, 2 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-gc.c b/builtin-gc.c\nindex 799c263..c6806d3 100644\n--- a/builtin-gc.c\n+++ b/builtin-gc.c\n@@ -23,7 +23,7 @@ static const char * const builtin_gc_usage[] = {\n };\n \n static int pack_refs = 1;\n-static int aggressive_window = -1;\n+static int aggressive_window = 250;\n static int gc_auto_threshold = 6700;\n static int gc_auto_pack_limit = 20;\n \n@@ -192,6 +192,7 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \n \tif (aggressive) {\n \t\tappend_option(argv_repack, \"-f\", MAX_ADD);\n+\t\tappend_option(argv_repack, \"--depth=250\", MAX_ADD);\n \t\tif (aggressive_window > 0) {\n \t\t\tsprintf(buf, \"--window=%d\", aggressive_window);\n \t\t\tappend_option(argv_repack, buf, MAX_ADD);\n-- \n1.5.3.7.2157.g9598e\n"},{"id":"62135","messageId":"200712061404.58827.ismail@pardus.org.tr","threadId":"11153","inReplyTo":"Pine.LNX.4.64.0712061150590.27959@racer.site","subject":"Re: Git and GCC","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-06T12:04:58Z","receivedAt":"2007-12-06T12:04:58Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Thursday 06 December 2007 13:57:06 Johannes Schindelin yazmıştı:\n[...]\n> So I fully expect an issue like Daniel's to be resolved in a matter of\n> minutes on the git list, if the OP gives us a chance.  If we are not even\n> Cc'ed, you are completely right, she or he probably does not want the\n> issue to be resolved.\n\nLets be fair about this, Ollie Wild already sent a mail about git-svn disk \nusage and there is no concrete solution yet, though it seems the bottleneck \nis known.\n\nRegards,\nismail\n\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"62144","messageId":"20071206134243.GA17037@thunk.org","threadId":"11153","inReplyTo":"Pine.LNX.4.64.0712061201580.27959@racer.site","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-12-06T13:42:44Z","receivedAt":"2007-12-06T13:42:44Z","isPatch":true,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Dec 06, 2007 at 12:03:38PM +0000, Johannes Schindelin wrote:\n> \n> The default was not to change the window or depth at all.  As suggested\n> by Jon Smirl, Linus Torvalds and others, default to\n> \n> \t--window=250 --depth=250\n\nI'd also suggest adding a comment in the man pages that this should\nonly be done rarely, and that it can potentially take a *long* time\n(i.e., overnight) for big repositories, and in general it's not worth\nthe effort to use --aggressive.\n\nApologies to Linus and to the gcc folks, since I was the one who\noriginally coded up gc --aggressive, and at the time my intent was\n\"rarely does it make sense, and it may take a long time\".  The reason\nwhy I didn't make the default --window and --depth larger is because\nat the time the biggest repo I had easy access to was the Linux\nkernel's, and there you rapidly hit diminishing returns at much\nsmaller numbers, so there was no real point in using --window=250\n--depth=250.\n\nLinus later pointed out that what we *really* should do is at some\npoint was to change repack -f to potentially retry to find a better\ndelta, but to reuse the existing delta if it was no worse.  That\nautomatically does the right thing in the case where you had\npreviously done a repack with --window=<large n> --depth=<large n>,\nbut then later try using \"gc --agressive\", which ends up doing a worse\njob and throwing away the information from the previous repack with\nlarge window and depth sizes.  Unfortunately no one ever got around to\nimplementing that.\n\nRegards,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"62146","messageId":"alpine.LFD.0.99999.0712060901120.555@xanadu.home","threadId":"11153","inReplyTo":"1196927361.13109.1.camel@brick","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T14:01:51Z","receivedAt":"2007-12-06T14:01:51Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 5 Dec 2007, Harvey Harrison wrote:\n\n> \n> > \tgit repack -a -d --depth=250 --window=250\n> > \n> \n> Since I have the whole gcc repo locally I'll give this a shot overnight\n> just to see what can be done at the extreme end or things.\n\nDon't forget to add -f as well.\n\n\nNicolas\n"},{"id":"62149","messageId":"alpine.LFD.0.99999.0712060908490.555@xanadu.home","threadId":"11153","inReplyTo":"20071206134243.GA17037@thunk.org","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T14:15:11Z","receivedAt":"2007-12-06T14:15:11Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Theodore Tso wrote:\n\n> Linus later pointed out that what we *really* should do is at some\n> point was to change repack -f to potentially retry to find a better\n> delta, but to reuse the existing delta if it was no worse.  That\n> automatically does the right thing in the case where you had\n> previously done a repack with --window=<large n> --depth=<large n>,\n> but then later try using \"gc --agressive\", which ends up doing a worse\n> job and throwing away the information from the previous repack with\n> large window and depth sizes.  Unfortunately no one ever got around to\n> implementing that.\n\nI did start looking at it, but there are subtle issues to consider, such \nas making sure not to create delta loops.  Currently this is avoided by \nnever involving already reused deltas in new delta chains, except for \nedge base objects.\n\nIOW, this requires some head scratching which I didn't have the time for \nso far.\n\n\nNicolas\n"},{"id":"62150","messageId":"alpine.LFD.0.99999.0712060915590.555@xanadu.home","threadId":"11153","inReplyTo":"20071206071503.GA19504@coredump.intra.peff.net","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T14:18:39Z","receivedAt":"2007-12-06T14:18:39Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Jeff King wrote:\n\n> On Thu, Dec 06, 2007 at 01:47:54AM -0500, Jon Smirl wrote:\n> \n> > The key to converting repositories of this size is RAM. 4GB minimum,\n> > more would be better. git-repack is not multi-threaded. There were a\n> > few attempts at making it multi-threaded but none were too successful.\n> > If I remember right, with loads of RAM, a repack on a 450MB repository\n> > was taking about five hours on a 2.8Ghz Core2. But this is something\n> > you only have to do once for the import. Later repacks will reuse the\n> > original deltas.\n> \n> Actually, Nicolas put quite a bit of work into multi-threading the\n> repack process; the results have been in master for some time, and will\n> be in the soon-to-be-released v1.5.4.\n> \n> The downside is that the threading partitions the object space, so the\n> resulting size is not necessarily as small (but I don't know that\n> anybody has done testing on large repos to find out how large the\n> difference is).\n\nQuick guesstimate is in the 1% ballpark.\n\n\nNicolas\n"},{"id":"62151","messageId":"20071206142254.GD5959@artemis.madism.org","threadId":"11153","inReplyTo":"Pine.LNX.4.64.0712061201580.27959@racer.site","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Pierre Habouzit","fromEmail":"madcoder@debian.org","sentAt":"2007-12-06T14:22:54Z","receivedAt":"2007-12-06T14:22:54Z","isPatch":true,"sender":{"key":"madcoder@debian.org","avatar":"https://avatars.githubusercontent.com/u/44708?v=4"},"body":"On Thu, Dec 06, 2007 at 12:03:38PM +0000, Johannes Schindelin wrote:\n> \n> The default was not to change the window or depth at all.  As suggested\n> by Jon Smirl, Linus Torvalds and others, default to\n> \n> \t--window=250 --depth=250\n\n  well, this will explode on many quite reasonnably sized systems. This\nshould also use a memory-limit that could be auto-guessed from the\nsystem total physical memory (50% of the actual memory could be a good\nidea e.g.).\n\n  On very large repositories, using that on the e.g. linux kernel, swaps\nlike hell on a machine with 1Go of ram, and almost nothing running on it\n(less than 200Mo of ram actually used)\n"},{"id":"62156","messageId":"1196955059.13633.3.camel@brick","threadId":"11153","inReplyTo":"Pine.LNX.4.64.0712061201580.27959@racer.site","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-06T15:30:59Z","receivedAt":"2007-12-06T15:30:59Z","isPatch":true,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"Wow\n\n/usr/bin/time git repack -a -d -f --window=250 --depth=250\n\n\n23266.37user 581.04system 7:41:25elapsed 86%CPU (0avgtext+0avgdata\n0maxresident)k\n0inputs+0outputs (419835major+123275804minor)pagefaults 0swaps\n\n-r--r--r-- 1 hharrison hharrison  29091872 2007-12-06 07:26\npack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.idx\n-r--r--r-- 1 hharrison hharrison 324094684 2007-12-06 07:26\npack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.pack\n\n\nThat extra delta depth really does make a difference.  Just over a\n300MB pack in the end, for all gcc branches/tags as of last night.\n\nCheers,\n\nHarvey\n"},{"id":"62159","messageId":"Pine.LNX.4.64.0712061552550.27959@racer.site","threadId":"11153","inReplyTo":"20071206142254.GD5959@artemis.madism.org","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-12-06T15:55:43Z","receivedAt":"2007-12-06T15:55:43Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Dec 2007, Pierre Habouzit wrote:\n\n> On Thu, Dec 06, 2007 at 12:03:38PM +0000, Johannes Schindelin wrote:\n> > \n> > The default was not to change the window or depth at all.  As \n> > suggested by Jon Smirl, Linus Torvalds and others, default to\n> > \n> > \t--window=250 --depth=250\n> \n>   well, this will explode on many quite reasonnably sized systems. This \n> should also use a memory-limit that could be auto-guessed from the \n> system total physical memory (50% of the actual memory could be a good \n> idea e.g.).\n> \n>   On very large repositories, using that on the e.g. linux kernel, swaps \n> like hell on a machine with 1Go of ram, and almost nothing running on it \n> (less than 200Mo of ram actually used)\n\nYes.\n\nHowever, I think that --aggressive should be aggressive, and if you decide \nto run it on a machine which lacks the muscle to be aggressive, well, you \nshould have known better.\n\nThe upside: if you run this on a strong machine and clone it to a weak \nmachine, you'll still have the benefit of a small pack (and you should \nmark it as .keep, too, to keep the benefit...)\n\nCiao,\nDscho\n"},{"id":"62160","messageId":"Pine.LNX.4.64.0712061555490.27959@racer.site","threadId":"11153","inReplyTo":"1196955059.13633.3.camel@brick","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-12-06T15:56:01Z","receivedAt":"2007-12-06T15:56:01Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Dec 2007, Harvey Harrison wrote:\n\n> -r--r--r-- 1 hharrison hharrison 324094684 2007-12-06 07:26\n> pack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.pack\n\nWow.\n\nCiao,\nDscho\n"},{"id":"62162","messageId":"alpine.LFD.0.9999.0712060803430.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"1196955059.13633.3.camel@brick","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T16:19:24Z","receivedAt":"2007-12-06T16:19:24Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Harvey Harrison wrote:\n> \n> 7:41:25elapsed 86%CPU\n\nHeh. And this is why you want to do it exactly *once*, and then just \nexport the end result for others ;)\n\n> -r--r--r-- 1 hharrison hharrison 324094684 2007-12-06 07:26 pack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.pack\n\nBut yeah, especially if you allow longer delta chains, the end result can \nbe much smaller (and what makes the one-time repack more expensive is the \nwindow size, not the delta chain - you could make the delta chains longer \nwith no cost overhead at packing time)\n\nHOWEVER. \n\nThe longer delta chains do make it potentially much more expensive to then \nuse old history. So there's a trade-off. And quite frankly, a delta depth \nof 250 is likely going to cause overflows in the delta cache (which is \nonly 256 entries in size *and* it's a hash, so it's going to start having \nhash conflicts long before hitting the 250 depth limit).\n\nSo when I said \"--depth=250 --window=250\", I chose those numbers more as \nan example of extremely aggressive packing, and I'm not at all sure that \nthe end result is necessarily wonderfully usable. It's going to save disk \nspace (and network bandwidth - the delta's will be re-used for the network \nprotocol too!), but there are definitely downsides too, and using long \ndelta chains may simply not be worth it in practice.\n\n(And some of it might just want to have git tuning, ie if people think \nthat long deltas are worth it, we could easily just expand on the delta \nhash, at the cost of some more memory used!)\n\nThat said, the good news is that working with *new* history will not be \naffected negatively, and if you want to be _really_ sneaky, there are ways \nto say \"create a pack that contains the history up to a version one year \nago, and be very aggressive about those old versions that we still want to \nhave around, but do a separate pack for newer stuff using less aggressive \nparameters\"\n\nSo this is something that can be tweaked, although we don't really have \nany really nice interfaces for stuff like that (ie the git delta cache \nsize is hardcoded in the sources and cannot be set in the config file, and \nthe \"pack old history more aggressively\" involves some manual scripting \nand knowing how \"git pack-objects\" works rather than any nice simple \ncommand line switch).\n\nSo the thing to take away from this is:\n - git is certainly flexible as hell\n - .. but to get the full power you may need to tweak things\n - .. happily you really only need to have one person to do the tweaking, \n   and the tweaked end results will be available to others that do not \n   need to know/care.\n\nAnd whether the difference between 320MB and 500MB is worth any really \ninvolved tweaking (considering the potential downsides), I really don't \nknow. Only testing will tell.\n\n\t\t\tLinus\n"},{"id":"62166","messageId":"8563zbpxde.fsf@lola.goethe.zz","threadId":"11153","inReplyTo":"Pine.LNX.4.64.0712061552550.27959@racer.site","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-12-06T17:05:17Z","receivedAt":"2007-12-06T17:05:17Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> However, I think that --aggressive should be aggressive, and if you\n> decide to run it on a machine which lacks the muscle to be aggressive,\n> well, you should have known better.\n\nThat's a rather cheap shot.  \"you should have known better\" than\nexpecting to be able to use a documented command and option because the\ngit developers happened to have a nicer machine...\n\n_How_ is one supposed to have known better?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"62169","messageId":"20071206173946.GA10845@sigill.intra.peff.net","threadId":"11153","inReplyTo":"alpine.LFD.0.99999.0712060915590.555@xanadu.home","subject":"Re: Git and GCC","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-06T17:39:47Z","receivedAt":"2007-12-06T17:39:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 06, 2007 at 09:18:39AM -0500, Nicolas Pitre wrote:\n\n> > The downside is that the threading partitions the object space, so the\n> > resulting size is not necessarily as small (but I don't know that\n> > anybody has done testing on large repos to find out how large the\n> > difference is).\n> \n> Quick guesstimate is in the 1% ballpark.\n\nFortunately, we now have numbers. Harvey Harrison reported repacking the\ngcc repo and getting these results:\n\n> /usr/bin/time git repack -a -d -f --window=250 --depth=250\n>\n> 23266.37user 581.04system 7:41:25elapsed 86%CPU (0avgtext+0avgdata 0maxresident)k\n> 0inputs+0outputs (419835major+123275804minor)pagefaults 0swaps\n>\n> -r--r--r-- 1 hharrison hharrison  29091872 2007-12-06 07:26 pack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.idx\n> -r--r--r-- 1 hharrison hharrison 324094684 2007-12-06 07:26 pack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.pack\n\nI tried the threaded repack with pack.threads = 3 on a dual-processor\nmachine, and got:\n\n  time git repack -a -d -f --window=250 --depth=250\n\n  real    309m59.849s\n  user    377m43.948s\n  sys     8m23.319s\n\n  -r--r--r-- 1 peff peff  28570088 2007-12-06 10:11 pack-1fa336f33126d762988ed6fc3f44ecbe0209da3c.idx\n  -r--r--r-- 1 peff peff 339922573 2007-12-06 10:11 pack-1fa336f33126d762988ed6fc3f44ecbe0209da3c.pack\n\nSo it is about 5% bigger. What is really disappointing is that we saved\nonly about 20% of the time. I didn't sit around watching the stages, but\nmy guess is that we spent a long time in the single threaded \"writing\nobjects\" stage with a thrashing delta cache.\n\n-Peff\n"},{"id":"62170","messageId":"alpine.LFD.0.99999.0712061246120.555@xanadu.home","threadId":"11153","inReplyTo":"20071206173946.GA10845@sigill.intra.peff.net","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T18:02:58Z","receivedAt":"2007-12-06T18:02:58Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Jeff King wrote:\n\n> On Thu, Dec 06, 2007 at 09:18:39AM -0500, Nicolas Pitre wrote:\n> \n> > > The downside is that the threading partitions the object space, so the\n> > > resulting size is not necessarily as small (but I don't know that\n> > > anybody has done testing on large repos to find out how large the\n> > > difference is).\n> > \n> > Quick guesstimate is in the 1% ballpark.\n> \n> Fortunately, we now have numbers. Harvey Harrison reported repacking the\n> gcc repo and getting these results:\n> \n> > /usr/bin/time git repack -a -d -f --window=250 --depth=250\n> >\n> > 23266.37user 581.04system 7:41:25elapsed 86%CPU (0avgtext+0avgdata 0maxresident)k\n> > 0inputs+0outputs (419835major+123275804minor)pagefaults 0swaps\n> >\n> > -r--r--r-- 1 hharrison hharrison  29091872 2007-12-06 07:26 pack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.idx\n> > -r--r--r-- 1 hharrison hharrison 324094684 2007-12-06 07:26 pack-1d46ca030c3d6d6b95ad316deb922be06b167a3d.pack\n> \n> I tried the threaded repack with pack.threads = 3 on a dual-processor\n> machine, and got:\n> \n>   time git repack -a -d -f --window=250 --depth=250\n> \n>   real    309m59.849s\n>   user    377m43.948s\n>   sys     8m23.319s\n> \n>   -r--r--r-- 1 peff peff  28570088 2007-12-06 10:11 pack-1fa336f33126d762988ed6fc3f44ecbe0209da3c.idx\n>   -r--r--r-- 1 peff peff 339922573 2007-12-06 10:11 pack-1fa336f33126d762988ed6fc3f44ecbe0209da3c.pack\n> \n> So it is about 5% bigger.\n\nRight.  I should probably revisit that idea of finding deltas across \npartition boundaries to mitigate that loss.  And those partitions could \nbe made coarser as well to reduce the number of such partition gaps \n(just increase the value of chunk_size on line 1648 in \nbuiltin-pack-objects.c).\n\n> What is really disappointing is that we saved\n> only about 20% of the time. I didn't sit around watching the stages, but\n> my guess is that we spent a long time in the single threaded \"writing\n> objects\" stage with a thrashing delta cache.\n\nMaybe you should run the non threaded repack on the same machine to have \na good comparison.  And if you have only 2 CPUs, you will have better \nperformances with pack.threads = 2, otherwise there'll be wasteful task \nswitching going on.\n\nAnd of course, if the delta cache is being trashed, that might be due to \nthe way the existing pack was previously packed.  Hence the current pack \nmight impact object _access_ when repacking them.  So for a really \nreally fair performance comparison, you'd have to preserve the original \npack and swap it back before each repack attempt.\n\n\nNicolas\n"},{"id":"62171","messageId":"4aca3dc20712061004g43f5902cw79bf633917d3ade9@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712052132450.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-06T18:04:35Z","receivedAt":"2007-12-06T18:04:35Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Dec 2007, Daniel Berlin wrote:\n> >\n> > Actually, it turns out that git-gc --aggressive does this dumb thing\n> > to pack files sometimes regardless of whether you converted from an\n> > SVN repo or not.\n>\n> Absolutely. git --aggressive is mostly dumb. It's really only useful for\n> the case of \"I know I have a *really* bad pack, and I want to throw away\n> all the bad packing decisions I have done\".\n>\n> To explain this, it's worth explaining (you are probably aware of it, but\n> let me go through the basics anyway) how git delta-chains work, and how\n> they are so different from most other systems.\n>\nI worked on Monotone and other systems that use object stores. for a\nlittle while :)\nIn particular, I believe GIT's original object store was based on\nMonotone, IIRC.\n\n> In other SCM's, a delta-chain is generally fixed. It might be \"forwards\"\n> or \"backwards\", and it might evolve a bit as you work with the repository,\n> but generally it's a chain of changes to a single file represented as some\n> kind of single SCM entity. In CVS, it's obviously the *,v file, and a lot\n> of other systems do rather similar things.\n\n>\n> Git also does delta-chains, but it does them a lot more \"loosely\". There\n> is no fixed entity. Delta's are generated against any random other version\n> that git deems to be a good delta candidate (with various fairly\n> successful heursitics), and there are absolutely no hard grouping rules.\n\nSure. SVN actually supports this (surprisingly), it just never happens\nto choose delta bases that aren't related by ancestry.  (IE it would\nhave absolutely no problem with you using random other parts of the\nrepository as delta bases, and i've played with it before).\n\nI actually advocated we move towards an object store model, as\nancestry can be a  crappy way of approximating similarity when you\nhave a lot of branches.\n\n> So the equivalent of \"git gc --aggressive\" - but done *properly* - is to\n> do (overnight) something like\n>\n>         git repack -a -d --depth=250 --window=250\n>\nI gave this a try overnight, and it definitely helps a lot.\nThanks!\n\n> And then it's going to take forever and a day (ie a \"do it overnight\"\n> thing). But the end result is that everybody downstream from that\n> repository will get much better packs, without having to spend any effort\n> on it themselves.\n>\n\nIf your forever and a day is spent figuring out which deltas to use,\nyou can reduce this significantly.\nIf it is spent writing out the data, it's much harder. :)\n"},{"id":"62174","messageId":"b609cb3b0712061024rc48022bhc3fbfba02061dd94@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712052132450.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"NightStrike","fromEmail":"nightstrike@gmail.com","sentAt":"2007-12-06T18:24:34Z","receivedAt":"2007-12-06T18:24:34Z","isPatch":false,"sender":{"key":"nightstrike@gmail.com","avatar":null},"body":"On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Dec 2007, Daniel Berlin wrote:\n> >\n> > Actually, it turns out that git-gc --aggressive does this dumb thing\n> > to pack files sometimes regardless of whether you converted from an\n> > SVN repo or not.\n> I'll send a patch to Junio to just remove the \"git gc --aggressive\"\n> documentation. It can be useful, but it generally is useful only when you\n> really understand at a very deep level what it's doing, and that\n> documentation doesn't help you do that.\n\nNo disrespect is meant by this reply.  I am just curious (and I am\nprobably misunderstanding something)..  Why remove all of the\ndocumentation entirely?  Wouldn't it be better to just document it\nmore thoroughly?  I thought you did a fine job in this post in\nexplaining its purpose, when to use it, when not to, etc.  Removing\nthe documention seems counter-intuitive when you've already gone to\nthe trouble of creating good documentation here in this post.\n"},{"id":"62175","messageId":"alpine.LFD.0.9999.0712061024090.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"4aca3dc20712061004g43f5902cw79bf633917d3ade9@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T18:29:43Z","receivedAt":"2007-12-06T18:29:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Daniel Berlin wrote:\n>\n> I worked on Monotone and other systems that use object stores. for a \n> little while :) In particular, I believe GIT's original object store was \n> based on Monotone, IIRC.\n\nYes and no. \n\nMonotone does what git does for the blobs. But there is a big difference \nin how git then does it for everything else too, ie trees and history. \nTree being in that object store in particular are very important, and one \nof the biggest deals for deltas (actually, for two reasons: most of the \ntime they don't change AT ALL if some subdirectory gets no changes and you \ndon't need any delta, and even when they do change, it's usually going to \ndelta very well, since it's usually just a small part that changes).\n\n> > And then it's going to take forever and a day (ie a \"do it overnight\"\n> > thing). But the end result is that everybody downstream from that\n> > repository will get much better packs, without having to spend any effort\n> > on it themselves.\n> \n> If your forever and a day is spent figuring out which deltas to use,\n> you can reduce this significantly.\n\nIt's almost all about figuring out the delta. Which is why *not* using \n\"-f\" (or \"--aggressive\") is such a big deal for normal operation, because \nthen you just skip it all.\n\n\t\tLinus\n"},{"id":"62176","messageId":"alpine.LFD.0.9999.0712061030560.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"20071206173946.GA10845@sigill.intra.peff.net","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T18:35:22Z","receivedAt":"2007-12-06T18:35:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Jeff King wrote:\n> \n> What is really disappointing is that we saved only about 20% of the \n> time. I didn't sit around watching the stages, but my guess is that we \n> spent a long time in the single threaded \"writing objects\" stage with a \n> thrashing delta cache.\n\nI don't think you spent all that much time writing the objects. That part \nisn't very intensive, it's mostly about the IO.\n\nI suspect you may simply be dominated by memory-throughput issues. The \ndelta matching doesn't cache all that well, and using two or more cores \nisn't going to help all that much if they are largely waiting for memory \n(and quite possibly also perhaps fighting each other for a shared cache? \nIs this a Core 2 with the shared L2?)\n\n\t\t\tLinus\n"},{"id":"62179","messageId":"alpine.LFD.0.9999.0712061036200.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"b609cb3b0712061024rc48022bhc3fbfba02061dd94@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T18:45:40Z","receivedAt":"2007-12-06T18:45:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, NightStrike wrote:\n> \n> No disrespect is meant by this reply.  I am just curious (and I am\n> probably misunderstanding something)..  Why remove all of the\n> documentation entirely?  Wouldn't it be better to just document it\n> more thoroughly?\n\nWell, part of it is that I don't think \"--aggressive\" as it is implemented \nright now is really almost *ever* the right answer. We could change the \nimplementation, of course, but generally the right thing to do is to not \nuse it (tweaking the \"--window\" and \"--depth\" manually for the repacking \nis likely the more natural thing to do).\n\nThe other part of the answer is that, when you *do* want to do what that \n\"--aggressive\" tries to achieve, it's such a special case event that while \nit should probably be documented, I don't think it should necessarily be \ndocumented where it is now (as part of \"git gc\"), but as part of a much \nmore technical manual for \"deep and subtle tricks you can play\".\n\n> I thought you did a fine job in this post in explaining its purpose, \n> when to use it, when not to, etc.  Removing the documention seems \n> counter-intuitive when you've already gone to the trouble of creating \n> good documentation here in this post.\n\nI'm so used to writing emails, and I *like* trying to explain what is \ngoing on, so I have no problems at all doing that kind of thing. However, \ntrying to write a manual or man-page or other technical documentation is \nsomething rather different.\n\nIOW, I like explaining git within the _context_ of a discussion or a \nparticular problem/issue. But documentation should work regardless of \ncontext (or at least set it up), and that's the part I am not so good at.\n\nIn other words, if somebody (hint hint) thinks my explanation was good and \nreadable, I'd love for them to try to turn it into real documentation by \nediting it up and creating enough context for it! But I'm nort personally \nvery likely to do that. I'd just send Junio the patch to remove a \nmisleading part of the documentation we have.\n\n\t\tLinus\n"},{"id":"62180","messageId":"9e4733910712061055p353775d8wd0321bc9c81297b7@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712061030560.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-06T18:55:55Z","receivedAt":"2007-12-06T18:55:55Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Dec 2007, Jeff King wrote:\n> >\n> > What is really disappointing is that we saved only about 20% of the\n> > time. I didn't sit around watching the stages, but my guess is that we\n> > spent a long time in the single threaded \"writing objects\" stage with a\n> > thrashing delta cache.\n>\n> I don't think you spent all that much time writing the objects. That part\n> isn't very intensive, it's mostly about the IO.\n>\n> I suspect you may simply be dominated by memory-throughput issues. The\n> delta matching doesn't cache all that well, and using two or more cores\n> isn't going to help all that much if they are largely waiting for memory\n> (and quite possibly also perhaps fighting each other for a shared cache?\n> Is this a Core 2 with the shared L2?)\n\nWhen I lasted looked at the code, the problem was in evenly dividing\nthe work. I was using a four core machine and most of the time one\ncore would end up with 3-5x the work of the lightest loaded core.\nSetting pack.threads up to 20 fixed the problem. With a high number of\nthreads I was able to get a 4hr pack to finished in something like\n1:15.\n\nA scheme where each core could work a minute without communicating to\nthe other cores would be best. It would also be more efficient if the\ncores could avoid having sync points between them.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62183","messageId":"alpine.LFD.0.99999.0712061403000.555@xanadu.home","threadId":"11153","inReplyTo":"9e4733910712061055p353775d8wd0321bc9c81297b7@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T19:08:12Z","receivedAt":"2007-12-06T19:08:12Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Jon Smirl wrote:\n\n> On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> >\n> > On Thu, 6 Dec 2007, Jeff King wrote:\n> > >\n> > > What is really disappointing is that we saved only about 20% of the\n> > > time. I didn't sit around watching the stages, but my guess is that we\n> > > spent a long time in the single threaded \"writing objects\" stage with a\n> > > thrashing delta cache.\n> >\n> > I don't think you spent all that much time writing the objects. That part\n> > isn't very intensive, it's mostly about the IO.\n> >\n> > I suspect you may simply be dominated by memory-throughput issues. The\n> > delta matching doesn't cache all that well, and using two or more cores\n> > isn't going to help all that much if they are largely waiting for memory\n> > (and quite possibly also perhaps fighting each other for a shared cache?\n> > Is this a Core 2 with the shared L2?)\n> \n> When I lasted looked at the code, the problem was in evenly dividing\n> the work. I was using a four core machine and most of the time one\n> core would end up with 3-5x the work of the lightest loaded core.\n> Setting pack.threads up to 20 fixed the problem. With a high number of\n> threads I was able to get a 4hr pack to finished in something like\n> 1:15.\n\nBut as far as I know you didn't try my latest incarnation which has been\navailable in Git's master branch for a few months already.\n\n\nNicolas\n"},{"id":"62186","messageId":"1196968371.18340.30.camel@ld0161-tx32","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712052132450.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Jon Loeliger","fromEmail":"jdl@freescale.com","sentAt":"2007-12-06T19:12:51Z","receivedAt":"2007-12-06T19:12:51Z","isPatch":false,"sender":{"key":"jdl@jdl.com","avatar":"https://gravatar.com/avatar/75ce9a10b151acd2c28ec4ab2136dba7b2ff1634530bd04b155981a749d08a64?d=mp&s=160"},"body":"On Thu, 2007-12-06 at 00:09, Linus Torvalds wrote:\n\n> Git also does delta-chains, but it does them a lot more \"loosely\". There \n> is no fixed entity. Delta's are generated against any random other version \n> that git deems to be a good delta candidate (with various fairly \n> successful heursitics), and there are absolutely no hard grouping rules.\n\nI'd like to learn more about that.  Can someone point me to\neither more documentation on it?  In the absence of that,\nperhaps a pointer to the source code that implements it?\n\nI guess one question I posit is, would it be more accurate\nto think of this as a \"delta net\" in a weighted graph rather\nthan a \"delta chain\"?\n\nThanks,\njdl\n"},{"id":"62190","messageId":"alpine.LFD.0.9999.0712061118050.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"1196968371.18340.30.camel@ld0161-tx32","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-06T19:39:49Z","receivedAt":"2007-12-06T19:39:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Jon Loeliger wrote:\n>\n> On Thu, 2007-12-06 at 00:09, Linus Torvalds wrote:\n> > Git also does delta-chains, but it does them a lot more \"loosely\". There \n> > is no fixed entity. Delta's are generated against any random other version \n> > that git deems to be a good delta candidate (with various fairly \n> > successful heursitics), and there are absolutely no hard grouping rules.\n> \n> I'd like to learn more about that.  Can someone point me to\n> either more documentation on it?  In the absence of that,\n> perhaps a pointer to the source code that implements it?\n\nWell, in a very real sense, what the delta code does is:\n - just list every single object in the whole repository\n - walk over each object, trying to find another object that it can be \n   written as a delta against\n - write out the result as a pack-file\n\nThat's simplified: we may not walk _all_ objects, for example: only a \nglobal repack does that (and most pack creations are actually for pushign \nand pulling between two repositories, so we only walk the objects that are \nin the source but not the destination repository).\n\nThe interesting phase is the \"walk each object, try to find a delta\" part. \nIn particular, you don't want to try to find a delta by comparing each \nobject to every other object out there (that would be O(n^2) in objects, \nand with a fairly high constant cost too!). So what it does is to sort the \nobjects by a few heuristics (type of object, base name that object was \nfound as when traversing a tree and size, and how recently it was found in \nthe history).\n\nAnd then over that sorted list, it tries to find deltas between entries \nthat are \"close\" to each other (and that's where the \"--window=xyz\" thing \ncomes in - it says how big the window is for objects being close. A \nsmaller window generates somewhat less good deltas, but takes a lot less \neffort to generate).\n\nThe source is in git/builtin-pack-objects.c, with the core of it being\n\n - try_delta() - try to generate a *single* delta when given an object \n   pair.\n\n - find_deltas() - do the actual list traversal\n\n - prepare_pack() and type_size_sort() - create the delta sort list from \n   the list of objects.\n\nbut that whole file is probably some of the more opaque parts of git.\n\n> I guess one question I posit is, would it be more accurate\n> to think of this as a \"delta net\" in a weighted graph rather\n> than a \"delta chain\"?\n\nIt's certainly not a simple chain, it's more of a set of acyclic directed \ngraphs in the object list. And yes, it's weigted by the size of the delta \nbetween objects, and the optimization problem is kind of akin to finding \nthe smallest spanning tree (well, forest - since you do *not* want to \ncreate one large graph, you also want to make the individual trees shallow \nenough that you don't have excessive delta depth).\n\nThere are good algorithms for finding minimum spanning trees, but this one \nis complicated by the fact that the biggest cost (by far!) is the \ncalculation of the weights itself. So rather than really worry about \nfinding the minimal tree/forest, the code needs to worry about not having \nto even calculate all the weights!\n\n(That, btw, is a common theme. A lot of git is about traversing graphs, \nlike the revision graph. And most of the trivial graph problems all assume \nthat you have the whole graph, but since the \"whole graph\" is the whole \nhistory of the repository, those algorithms are totally worthless, since \nthey are fundamentally much too expensive - if we have to generate the \nwhole history, we're already screwed for a big project. So things like \nrevision graph calculation, the main performance issue is to avoid having \nto even *look* at parts of the graph that we don't need to see!)\n\n\t\t\tLinus\n"},{"id":"62192","messageId":"7vk5nrd1yq.fsf@gitster.siamese.dyndns.org","threadId":"11153","inReplyTo":"1196968371.18340.30.camel@ld0161-tx32","subject":"Re: Git and GCC","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-06T20:04:29Z","receivedAt":"2007-12-06T20:04:29Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jon Loeliger <jdl@freescale.com> writes:\n\n> On Thu, 2007-12-06 at 00:09, Linus Torvalds wrote:\n>\n>> Git also does delta-chains, but it does them a lot more \"loosely\". There \n>> is no fixed entity. Delta's are generated against any random other version \n>> that git deems to be a good delta candidate (with various fairly \n>> successful heursitics), and there are absolutely no hard grouping rules.\n>\n> I'd like to learn more about that.  Can someone point me to\n> either more documentation on it?  In the absence of that,\n> perhaps a pointer to the source code that implements it?\n\nSee Documentation/technical/pack-heuristics.txt,\nbut the document predates and does not talk about delta\nreusing, which was covered here:\n\n    http://thread.gmane.org/gmane.comp.version-control.git/16223/focus=16267\n\n> I guess one question I posit is, would it be more accurate\n> to think of this as a \"delta net\" in a weighted graph rather\n> than a \"delta chain\"?\n\nYes.\n"},{"id":"62199","messageId":"7vabonczad.fsf@gitster.siamese.dyndns.org","threadId":"11153","inReplyTo":"7vk5nrd1yq.fsf@gitster.siamese.dyndns.org","subject":"Re: Git and GCC","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-06T21:02:18Z","receivedAt":"2007-12-06T21:02:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Jon Loeliger <jdl@freescale.com> writes:\n>\n>> I'd like to learn more about that.  Can someone point me to\n>> either more documentation on it?  In the absence of that,\n>> perhaps a pointer to the source code that implements it?\n>\n> See Documentation/technical/pack-heuristics.txt,\n\nA somewhat funny thing about this is ...\n\n$ git show --stat --summary b116b297\ncommit b116b297a80b54632256eb89dd22ea2b140de622\nAuthor: Jon Loeliger <jdl@jdl.com>\nDate:   Thu Mar 2 19:19:29 2006 -0600\n\n    Added Packing Heursitics IRC writeup.\n    \n    Signed-off-by: Jon Loeliger <jdl@jdl.com>\n    Signed-off-by: Junio C Hamano <junkio@cox.net>\n\n Documentation/technical/pack-heuristics.txt |  466 +++++++++++++++++++++++++++\n 1 files changed, 466 insertions(+), 0 deletions(-)\n create mode 100644 Documentation/technical/pack-heuristics.txt\n"},{"id":"62202","messageId":"9e4733910712061339n3aef023r22e5b73aac120c8a@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.99999.0712061403000.555@xanadu.home","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-06T21:39:02Z","receivedAt":"2007-12-06T21:39:02Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > When I lasted looked at the code, the problem was in evenly dividing\n> > the work. I was using a four core machine and most of the time one\n> > core would end up with 3-5x the work of the lightest loaded core.\n> > Setting pack.threads up to 20 fixed the problem. With a high number of\n> > threads I was able to get a 4hr pack to finished in something like\n> > 1:15.\n>\n> But as far as I know you didn't try my latest incarnation which has been\n> available in Git's master branch for a few months already.\n\nI've deleted all my giant packs. Using the kernel pack:\n4GB Q6600\n\nUsing the current thread pack code I get these results.\n\nThe interesting case is the last one. I set it to 15 threads and\nmonitored with 'top'.\nFor 0-60% compression I was at 300% CPU, 60-74% was 200% CPU and\n74-100% was 100% CPU. It never used all for cores. The only other\nthings running were top and my desktop. This is the same load\nbalancing problem I observed earlier. Much more clock time was spent\nin the 2/1 core phases than the 3 core one.\n\nThreaded, threads = 5\n\njonsmirl@terra:/home/linux$ time git repack -a -d -f\nCounting objects: 648366, done.\nCompressing objects: 100% (647457/647457), done.\nWriting objects: 100% (648366/648366), done.\nTotal 648366 (delta 528994), reused 0 (delta 0)\n\nreal    1m31.395s\nuser    2m59.239s\nsys     0m3.048s\njonsmirl@terra:/home/linux$\n\n12 seconds counting\n53 seconds compressing\n38 seconds writing\n\nWithout threads,\n\njonsmirl@terra:/home/linux$ time git repack -a -d -f\nwarning: no threads support, ignoring pack.threads\nCounting objects: 648366, done.\nCompressing objects: 100% (647457/647457), done.\nWriting objects: 100% (648366/648366), done.\nTotal 648366 (delta 528999), reused 0 (delta 0)\n\nreal    2m54.849s\nuser    2m51.267s\nsys     0m1.412s\njonsmirl@terra:/home/linux$\n\nThreaded, threads = 5\n\njonsmirl@terra:/home/linux$ time git repack -a -d -f --depth=250 --window=250\nCounting objects: 648366, done.\nCompressing objects: 100% (647457/647457), done.\nWriting objects: 100% (648366/648366), done.\nTotal 648366 (delta 539080), reused 0 (delta 0)\n\nreal    9m18.032s\nuser    19m7.484s\nsys     0m3.880s\njonsmirl@terra:/home/linux$\n\njonsmirl@terra:/home/linux/.git/objects/pack$ ls -l\ntotal 182156\n-r--r--r-- 1 jonsmirl jonsmirl  15561848 2007-12-06 16:15\npack-f1f8637d2c68eb1c964ec7c1877196c0c7513412.idx\n-r--r--r-- 1 jonsmirl jonsmirl 170768761 2007-12-06 16:15\npack-f1f8637d2c68eb1c964ec7c1877196c0c7513412.pack\njonsmirl@terra:/home/linux/.git/objects/pack$\n\nNon-threaded:\n\njonsmirl@terra:/home/linux$ time git repack -a -d -f --depth=250 --window=250\nwarning: no threads support, ignoring pack.threads\nCounting objects: 648366, done.\nCompressing objects: 100% (647457/647457), done.\nWriting objects: 100% (648366/648366), done.\nTotal 648366 (delta 539080), reused 0 (delta 0)\n\nreal    18m51.183s\nuser    18m46.538s\nsys     0m1.604s\njonsmirl@terra:/home/linux$\n\n\njonsmirl@terra:/home/linux/.git/objects/pack$ ls -l\ntotal 182156\n-r--r--r-- 1 jonsmirl jonsmirl  15561848 2007-12-06 15:33\npack-f1f8637d2c68eb1c964ec7c1877196c0c7513412.idx\n-r--r--r-- 1 jonsmirl jonsmirl 170768761 2007-12-06 15:33\npack-f1f8637d2c68eb1c964ec7c1877196c0c7513412.pack\njonsmirl@terra:/home/linux/.git/objects/pack$\n\nThreaded, threads = 15\n\njonsmirl@terra:/home/linux$ time git repack -a -d -f --depth=250 --window=250\nCounting objects: 648366, done.\nCompressing objects: 100% (647457/647457), done.\nWriting objects: 100% (648366/648366), done.\nTotal 648366 (delta 539080), reused 0 (delta 0)\n\nreal    9m18.325s\nuser    19m14.340s\nsys     0m3.996s\njonsmirl@terra:/home/linux$\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62205","messageId":"alpine.LFD.0.99999.0712061645120.555@xanadu.home","threadId":"11153","inReplyTo":"9e4733910712061339n3aef023r22e5b73aac120c8a@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T22:08:03Z","receivedAt":"2007-12-06T22:08:03Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Jon Smirl wrote:\n\n> On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > > When I lasted looked at the code, the problem was in evenly dividing\n> > > the work. I was using a four core machine and most of the time one\n> > > core would end up with 3-5x the work of the lightest loaded core.\n> > > Setting pack.threads up to 20 fixed the problem. With a high number of\n> > > threads I was able to get a 4hr pack to finished in something like\n> > > 1:15.\n> >\n> > But as far as I know you didn't try my latest incarnation which has been\n> > available in Git's master branch for a few months already.\n> \n> I've deleted all my giant packs. Using the kernel pack:\n> 4GB Q6600\n> \n> Using the current thread pack code I get these results.\n> \n> The interesting case is the last one. I set it to 15 threads and\n> monitored with 'top'.\n> For 0-60% compression I was at 300% CPU, 60-74% was 200% CPU and\n> 74-100% was 100% CPU. It never used all for cores. The only other\n> things running were top and my desktop. This is the same load\n> balancing problem I observed earlier.\n\nWell, that's possible with a window 25 times larger than the default.\n\nThe load balancing is solved with a master thread serving relatively \nsmall object list segments to any work thread that finished with its \nprevious segment.  But the size for those segments is currently fixed to \nwindow * 1000 which is way too large when window == 250.\n\nI have to find a way to auto-tune that segment size somehow.\n\nBut with the default window size there should not be any such noticeable \nload balancing problem.\n\nNote that threading only happens in the compression phase.  The count \nand write phase are hardly paralleled.\n\n\nNicolas\n"},{"id":"62207","messageId":"9e4733910712061411y77f800dcx46bb8fdd5d97941f@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.99999.0712061645120.555@xanadu.home","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-06T22:11:48Z","receivedAt":"2007-12-06T22:11:48Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Thu, 6 Dec 2007, Jon Smirl wrote:\n>\n> > On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > > > When I lasted looked at the code, the problem was in evenly dividing\n> > > > the work. I was using a four core machine and most of the time one\n> > > > core would end up with 3-5x the work of the lightest loaded core.\n> > > > Setting pack.threads up to 20 fixed the problem. With a high number of\n> > > > threads I was able to get a 4hr pack to finished in something like\n> > > > 1:15.\n> > >\n> > > But as far as I know you didn't try my latest incarnation which has been\n> > > available in Git's master branch for a few months already.\n> >\n> > I've deleted all my giant packs. Using the kernel pack:\n> > 4GB Q6600\n> >\n> > Using the current thread pack code I get these results.\n> >\n> > The interesting case is the last one. I set it to 15 threads and\n> > monitored with 'top'.\n> > For 0-60% compression I was at 300% CPU, 60-74% was 200% CPU and\n> > 74-100% was 100% CPU. It never used all for cores. The only other\n> > things running were top and my desktop. This is the same load\n> > balancing problem I observed earlier.\n>\n> Well, that's possible with a window 25 times larger than the default.\n>\n> The load balancing is solved with a master thread serving relatively\n> small object list segments to any work thread that finished with its\n> previous segment.  But the size for those segments is currently fixed to\n> window * 1000 which is way too large when window == 250.\n>\n> I have to find a way to auto-tune that segment size somehow.\n\nThat would be nice. Threading is most important on the giant\npack/window combinations. The normal case is fast enough that I don't\nreal notice it. These giant pack/window combos can run 8-10 hours.\n\n>\n> But with the default window size there should not be any such noticeable\n> load balancing problem.\n\nI only spend 30 seconds in the compression phase without making the\nwindow larger. It's not long enough to really see what is going on.\n\n>\n> Note that threading only happens in the compression phase.  The count\n> and write phase are hardly paralleled.\n>\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62208","messageId":"9e4733910712061422w139273c0gf3cfb04c6ba8c509@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.99999.0712061645120.555@xanadu.home","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-06T22:22:04Z","receivedAt":"2007-12-06T22:22:04Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Thu, 6 Dec 2007, Jon Smirl wrote:\n>\n> > On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > > > When I lasted looked at the code, the problem was in evenly dividing\n> > > > the work. I was using a four core machine and most of the time one\n> > > > core would end up with 3-5x the work of the lightest loaded core.\n> > > > Setting pack.threads up to 20 fixed the problem. With a high number of\n> > > > threads I was able to get a 4hr pack to finished in something like\n> > > > 1:15.\n> > >\n> > > But as far as I know you didn't try my latest incarnation which has been\n> > > available in Git's master branch for a few months already.\n> >\n> > I've deleted all my giant packs. Using the kernel pack:\n> > 4GB Q6600\n> >\n> > Using the current thread pack code I get these results.\n> >\n> > The interesting case is the last one. I set it to 15 threads and\n> > monitored with 'top'.\n> > For 0-60% compression I was at 300% CPU, 60-74% was 200% CPU and\n> > 74-100% was 100% CPU. It never used all for cores. The only other\n> > things running were top and my desktop. This is the same load\n> > balancing problem I observed earlier.\n>\n> Well, that's possible with a window 25 times larger than the default.\n\nWhy did it never use more than three cores?\n\n>\n> The load balancing is solved with a master thread serving relatively\n> small object list segments to any work thread that finished with its\n> previous segment.  But the size for those segments is currently fixed to\n> window * 1000 which is way too large when window == 250.\n>\n> I have to find a way to auto-tune that segment size somehow.\n>\n> But with the default window size there should not be any such noticeable\n> load balancing problem.\n>\n> Note that threading only happens in the compression phase.  The count\n> and write phase are hardly paralleled.\n>\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62209","messageId":"85r6hzo3y8.fsf@lola.goethe.zz","threadId":"11153","inReplyTo":"7vabonczad.fsf@gitster.siamese.dyndns.org","subject":"Re: Git and GCC","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-12-06T22:26:07Z","receivedAt":"2007-12-06T22:26:07Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> Jon Loeliger <jdl@freescale.com> writes:\n>>\n>>> I'd like to learn more about that.  Can someone point me to\n>>> either more documentation on it?  In the absence of that,\n>>> perhaps a pointer to the source code that implements it?\n>>\n>> See Documentation/technical/pack-heuristics.txt,\n>\n> A somewhat funny thing about this is ...\n>\n> $ git show --stat --summary b116b297\n> commit b116b297a80b54632256eb89dd22ea2b140de622\n> Author: Jon Loeliger <jdl@jdl.com>\n> Date:   Thu Mar 2 19:19:29 2006 -0600\n>\n>     Added Packing Heursitics IRC writeup.\n\nAh, fishing for compliments.  The cookie baking season...\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"62210","messageId":"alpine.LFD.0.99999.0712061726240.555@xanadu.home","threadId":"11153","inReplyTo":"9e4733910712061422w139273c0gf3cfb04c6ba8c509@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-06T22:30:44Z","receivedAt":"2007-12-06T22:30:44Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Jon Smirl wrote:\n\n> On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Thu, 6 Dec 2007, Jon Smirl wrote:\n> >\n> > > On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > > > > When I lasted looked at the code, the problem was in evenly dividing\n> > > > > the work. I was using a four core machine and most of the time one\n> > > > > core would end up with 3-5x the work of the lightest loaded core.\n> > > > > Setting pack.threads up to 20 fixed the problem. With a high number of\n> > > > > threads I was able to get a 4hr pack to finished in something like\n> > > > > 1:15.\n> > > >\n> > > > But as far as I know you didn't try my latest incarnation which has been\n> > > > available in Git's master branch for a few months already.\n> > >\n> > > I've deleted all my giant packs. Using the kernel pack:\n> > > 4GB Q6600\n> > >\n> > > Using the current thread pack code I get these results.\n> > >\n> > > The interesting case is the last one. I set it to 15 threads and\n> > > monitored with 'top'.\n> > > For 0-60% compression I was at 300% CPU, 60-74% was 200% CPU and\n> > > 74-100% was 100% CPU. It never used all for cores. The only other\n> > > things running were top and my desktop. This is the same load\n> > > balancing problem I observed earlier.\n> >\n> > Well, that's possible with a window 25 times larger than the default.\n> \n> Why did it never use more than three cores?\n\nYou have 648366 objects total, and only 647457 of them are subject to \ndelta compression.\n\nWith a window size of 250 and a default thread segment of window * 1000 \nthat means only 3 segments will be distributed to threads, hence only 3 \nthreads with work to do.\n\n\nNicolas\n"},{"id":"62211","messageId":"20071206143827.004991f8.rdunlap@xenotime.net","threadId":"11153","inReplyTo":"85r6hzo3y8.fsf@lola.goethe.zz","subject":"[OT] Re: Git and GCC","fromName":"Randy Dunlap","fromEmail":"rdunlap@xenotime.net","sentAt":"2007-12-06T22:38:27Z","receivedAt":"2007-12-06T22:38:27Z","isPatch":false,"sender":{"key":"rdunlap@xenotime.net","avatar":null},"body":"On Thu, 06 Dec 2007 23:26:07 +0100 David Kastrup wrote:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > Junio C Hamano <gitster@pobox.com> writes:\n> >\n> >> Jon Loeliger <jdl@freescale.com> writes:\n> >>\n> >>> I'd like to learn more about that.  Can someone point me to\n> >>> either more documentation on it?  In the absence of that,\n> >>> perhaps a pointer to the source code that implements it?\n> >>\n> >> See Documentation/technical/pack-heuristics.txt,\n> >\n> > A somewhat funny thing about this is ...\n> >\n> > $ git show --stat --summary b116b297\n> > commit b116b297a80b54632256eb89dd22ea2b140de622\n> > Author: Jon Loeliger <jdl@jdl.com>\n> > Date:   Thu Mar 2 19:19:29 2006 -0600\n> >\n> >     Added Packing Heursitics IRC writeup.\n> \n> Ah, fishing for compliments.  The cookie baking season...\n\nIndeed.  Here are some really good & sweet recipes (IMHO).\n\nhttp://www.xenotime.net/linux/recipes/\n\n\n---\n~Randy\nFeatures and documentation: http://lwn.net/Articles/260136/\n"},{"id":"62212","messageId":"9e4733910712061444i64c115e2y94f6212dd7a4ddda@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.99999.0712061726240.555@xanadu.home","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-06T22:44:33Z","receivedAt":"2007-12-06T22:44:33Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Nicolas Pitre <nico@cam.org> wrote:\n> > > Well, that's possible with a window 25 times larger than the default.\n> >\n> > Why did it never use more than three cores?\n>\n> You have 648366 objects total, and only 647457 of them are subject to\n> delta compression.\n>\n> With a window size of 250 and a default thread segment of window * 1000\n> that means only 3 segments will be distributed to threads, hence only 3\n> threads with work to do.\n\nOne little tweak and the clock time drops from 9.5 to 6 minutes. The\ntweak makes all four cores work.\n\njonsmirl@terra:/home/apps/git$ git diff\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 4f44658..e0dd12e 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -1645,7 +1645,7 @@ static void ll_find_deltas(struct object_entry\n**list, unsigned list_size,\n        }\n\n        /* this should be auto-tuned somehow */\n-       chunk_size = window * 1000;\n+       chunk_size = window * 50;\n\n        do {\n                unsigned sublist_size = chunk_size;\n\n\njonsmirl@terra:/home/linux/.git$ time git repack -a -d -f --depth=250\n--window=250\nCounting objects: 648366, done.\nCompressing objects: 100% (647457/647457), done.\nWriting objects: 100% (648366/648366), done.\nTotal 648366 (delta 539043), reused 0 (delta 0)\n\nreal    6m2.109s\nuser    20m0.491s\nsys     0m4.608s\njonsmirl@terra:/home/linux/.git$\n\n\n\n>\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62220","messageId":"m3y7c7tkbs.fsf@roke.D-201","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712061118050.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-07T00:29:56Z","receivedAt":"2007-12-07T00:29:56Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Thu, 6 Dec 2007, Jon Loeliger wrote:\n\n>> I guess one question I posit is, would it be more accurate\n>> to think of this as a \"delta net\" in a weighted graph rather\n>> than a \"delta chain\"?\n> \n> It's certainly not a simple chain, it's more of a set of acyclic directed \n> graphs in the object list. And yes, it's weigted by the size of the delta \n> between objects, and the optimization problem is kind of akin to finding \n> the smallest spanning tree (well, forest - since you do *not* want to \n> create one large graph, you also want to make the individual trees shallow \n> enough that you don't have excessive delta depth).\n> \n> There are good algorithms for finding minimum spanning trees, but this one \n> is complicated by the fact that the biggest cost (by far!) is the \n> calculation of the weights itself. So rather than really worry about \n> finding the minimal tree/forest, the code needs to worry about not having \n> to even calculate all the weights!\n> \n> (That, btw, is a common theme. A lot of git is about traversing graphs, \n> like the revision graph. And most of the trivial graph problems all assume \n> that you have the whole graph, but since the \"whole graph\" is the whole \n> history of the repository, those algorithms are totally worthless, since \n> they are fundamentally much too expensive - if we have to generate the \n> whole history, we're already screwed for a big project. So things like \n> revision graph calculation, the main performance issue is to avoid having \n> to even *look* at parts of the graph that we don't need to see!)\n\nHmmm...\n\nI think that these two problems (find minimal spanning forest with\nlimited depth and traverse graph) with the additional constraint to\navoid calculating weights / avoid calculating whole graph would be\na good problem to present at CompSci course.\n\nJust a thought...\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"62229","messageId":"1196995353.22471.20.camel@brick","threadId":"11153","inReplyTo":"4aca3dc20712061004g43f5902cw79bf633917d3ade9@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-07T02:42:33Z","receivedAt":"2007-12-07T02:42:33Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Thu, 2007-12-06 at 13:04 -0500, Daniel Berlin wrote:\n> On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> \n> > So the equivalent of \"git gc --aggressive\" - but done *properly* - is to\n> > do (overnight) something like\n> >\n> >         git repack -a -d --depth=250 --window=250\n> >\n> I gave this a try overnight, and it definitely helps a lot.\n> Thanks!\n\nI've updated the public mirror repo with the very-packed version.\n\nPeople cloning it now should get the just over 300MB repo now.\n\ngit.infradead.org/gcc.git\n\n\nCheers,\n\nHarvey\n"},{"id":"62231","messageId":"alpine.LFD.0.9999.0712061857060.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"1196995353.22471.20.camel@brick","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-07T03:01:57Z","receivedAt":"2007-12-07T03:01:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Harvey Harrison wrote:\n> \n> I've updated the public mirror repo with the very-packed version.\n\nSide note: it might be interesting to compare timings for \nhistory-intensive stuff with and without this kind of very-packed \nsituation.\n\nThe very density of a smaller pack-file might be enough to overcome the \ndownsides (more CPU time to apply longer delta-chains), but regardless, \nreal numbers talks, bullshit walks. So wouldn't it be nice to have real \nnumbers?\n\nOne easy way to get real numbers for history would be to just time some \nreasonably costly operation that uses lots of history. Ie just do a \n\n\ttime git blame -C gcc/regclass.c > /dev/null\n\nand see if the deeper delta chains are very expensive.\n\n(Yeah, the above is pretty much designed to be the worst possible case for \nthis kind of aggressive history packing, but I don't know if that choice \nof file to try to annotate is a good choice or not. I suspect that \"git \nblame -C\" with a CVS import is just horrid, because CVS commits tend to be \npretty big and nasty and not as localized as we've tried to make things in \nthe kernel, so doing the code copy detection is probably horrendously \nexpensive)\n\n\t\t\tLinus\n"},{"id":"62233","messageId":"20071206.193121.40404287.davem@davemloft.net","threadId":"11153","inReplyTo":"20071206173946.GA10845@sigill.intra.peff.net","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-07T03:31:21Z","receivedAt":"2007-12-07T03:31:21Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: Jeff King <peff@peff.net>\nDate: Thu, 6 Dec 2007 12:39:47 -0500\n\n> I tried the threaded repack with pack.threads = 3 on a dual-processor\n> machine, and got:\n> \n>   time git repack -a -d -f --window=250 --depth=250\n> \n>   real    309m59.849s\n>   user    377m43.948s\n>   sys     8m23.319s\n> \n>   -r--r--r-- 1 peff peff  28570088 2007-12-06 10:11 pack-1fa336f33126d762988ed6fc3f44ecbe0209da3c.idx\n>   -r--r--r-- 1 peff peff 339922573 2007-12-06 10:11 pack-1fa336f33126d762988ed6fc3f44ecbe0209da3c.pack\n> \n> So it is about 5% bigger. What is really disappointing is that we saved\n> only about 20% of the time. I didn't sit around watching the stages, but\n> my guess is that we spent a long time in the single threaded \"writing\n> objects\" stage with a thrashing delta cache.\n\nIf someone can give me a good way to run this test case I can\nhave my 64-cpu Niagara-2 box crunch on this and see how fast\nit goes and how much larger the resulting pack file is.\n"},{"id":"62236","messageId":"9e4733910712062006l651571f3w7f76ce64c6650dff@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712061857060.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-07T04:06:07Z","receivedAt":"2007-12-07T04:06:07Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Dec 2007, Harvey Harrison wrote:\n> >\n> > I've updated the public mirror repo with the very-packed version.\n>\n> Side note: it might be interesting to compare timings for\n> history-intensive stuff with and without this kind of very-packed\n> situation.\n>\n> The very density of a smaller pack-file might be enough to overcome the\n> downsides (more CPU time to apply longer delta-chains), but regardless,\n> real numbers talks, bullshit walks. So wouldn't it be nice to have real\n> numbers?\n>\n> One easy way to get real numbers for history would be to just time some\n> reasonably costly operation that uses lots of history. Ie just do a\n>\n>         time git blame -C gcc/regclass.c > /dev/null\n>\n> and see if the deeper delta chains are very expensive.\n\njonsmirl@terra:/video/gcc$ time git blame -C gcc/regclass.c > /dev/null\n\nreal    1m21.967s\nuser    1m21.329s\nsys     0m0.640s\n\nThe Mozilla repo is at least 50% larger than the gcc one. It took me\n23 minutes to repack the gcc one on my $800 Dell. The trick to this is\nlots of RAM and 64b. There is little disk IO during the compression\nphase, everything is cached.\n\nI have a 4.8GB git process with 4GB of physical memory. Everything\nstarted slowing down a lot when the process got that big. Does git\nreally need 4.8GB to repack? I could only keep 3.4GB resident. Luckily\nthis happen at 95% completion. With 8GB of memory you should be able\nto do this repack in under 20 minutes.\n\njonsmirl@terra:/video/gcc$ time git repack -a -d -f --depth=250 --window=250\nreal    22m54.380s\nuser    69m18.948s\nsys     0m23.773s\n\n\n> (Yeah, the above is pretty much designed to be the worst possible case for\n> this kind of aggressive history packing, but I don't know if that choice\n> of file to try to annotate is a good choice or not. I suspect that \"git\n> blame -C\" with a CVS import is just horrid, because CVS commits tend to be\n> pretty big and nasty and not as localized as we've tried to make things in\n> the kernel, so doing the code copy detection is probably horrendously\n> expensive)\n>\n>                         Linus\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62237","messageId":"alpine.LFD.0.99999.0712062315000.555@xanadu.home","threadId":"11153","inReplyTo":"9e4733910712062006l651571f3w7f76ce64c6650dff@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-07T04:21:32Z","receivedAt":"2007-12-07T04:21:32Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 6 Dec 2007, Jon Smirl wrote:\n\n> I have a 4.8GB git process with 4GB of physical memory. Everything\n> started slowing down a lot when the process got that big. Does git\n> really need 4.8GB to repack? I could only keep 3.4GB resident. Luckily\n> this happen at 95% completion. With 8GB of memory you should be able\n> to do this repack in under 20 minutes.\n\nProbably you have too many cached delta results.  By default, every \ndelta smaller than 1000 bytes is kept in memory until the write phase.  \nTry using pack.deltacachesize = 256M or lower, or try disabling this \ncaching entirely with pack.deltacachelimit = 0.\n\n\nNicolas\n"},{"id":"62239","messageId":"alpine.LFD.0.9999.0712062120100.13796@woody.linux-foundation.org","threadId":"11153","inReplyTo":"9e4733910712062006l651571f3w7f76ce64c6650dff@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-07T05:21:35Z","receivedAt":"2007-12-07T05:21:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Dec 2007, Jon Smirl wrote:\n> >\n> >         time git blame -C gcc/regclass.c > /dev/null\n> \n> jonsmirl@terra:/video/gcc$ time git blame -C gcc/regclass.c > /dev/null\n> \n> real    1m21.967s\n> user    1m21.329s\n\nWell, I was also hoping for a \"compared to not-so-aggressive packing\" \nnumber on the same machine.. IOW, what I was wondering is whether there is \na visible performance downside to the deeper delta chains in the 300MB \npack vs the (less aggressive) 500MB pack.\n\n\t\tLinus\n"},{"id":"62242","messageId":"b609cb3b0712062136y154ec1c3g875ba33e470355a8@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712061036200.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"NightStrike","fromEmail":"nightstrike@gmail.com","sentAt":"2007-12-07T05:36:59Z","receivedAt":"2007-12-07T05:36:59Z","isPatch":false,"sender":{"key":"nightstrike@gmail.com","avatar":null},"body":"On 12/6/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Dec 2007, NightStrike wrote:\n> >\n> > No disrespect is meant by this reply.  I am just curious (and I am\n> > probably misunderstanding something)..  Why remove all of the\n> > documentation entirely?  Wouldn't it be better to just document it\n> > more thoroughly?\n>\n> Well, part of it is that I don't think \"--aggressive\" as it is implemented\n> right now is really almost *ever* the right answer. We could change the\n> implementation, of course, but generally the right thing to do is to not\n> use it (tweaking the \"--window\" and \"--depth\" manually for the repacking\n> is likely the more natural thing to do).\n>\n> The other part of the answer is that, when you *do* want to do what that\n> \"--aggressive\" tries to achieve, it's such a special case event that while\n> it should probably be documented, I don't think it should necessarily be\n> documented where it is now (as part of \"git gc\"), but as part of a much\n> more technical manual for \"deep and subtle tricks you can play\".\n>\n> > I thought you did a fine job in this post in explaining its purpose,\n> > when to use it, when not to, etc.  Removing the documention seems\n> > counter-intuitive when you've already gone to the trouble of creating\n> > good documentation here in this post.\n>\n> I'm so used to writing emails, and I *like* trying to explain what is\n> going on, so I have no problems at all doing that kind of thing. However,\n> trying to write a manual or man-page or other technical documentation is\n> something rather different.\n>\n> IOW, I like explaining git within the _context_ of a discussion or a\n> particular problem/issue. But documentation should work regardless of\n> context (or at least set it up), and that's the part I am not so good at.\n>\n> In other words, if somebody (hint hint) thinks my explanation was good and\n> readable, I'd love for them to try to turn it into real documentation by\n> editing it up and creating enough context for it! But I'm nort personally\n> very likely to do that. I'd just send Junio the patch to remove a\n> misleading part of the documentation we have.\n\nhehe.. I'd love to, actually.  I can work on it next week.\n"},{"id":"62246","messageId":"20071207063848.GA13101@coredump.intra.peff.net","threadId":"11153","inReplyTo":"20071206.193121.40404287.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-07T06:38:48Z","receivedAt":"2007-12-07T06:38:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 06, 2007 at 07:31:21PM -0800, David Miller wrote:\n\n> > So it is about 5% bigger. What is really disappointing is that we saved\n> > only about 20% of the time. I didn't sit around watching the stages, but\n> > my guess is that we spent a long time in the single threaded \"writing\n> > objects\" stage with a thrashing delta cache.\n> \n> If someone can give me a good way to run this test case I can\n> have my 64-cpu Niagara-2 box crunch on this and see how fast\n> it goes and how much larger the resulting pack file is.\n\nThat would be fun to see. The procedure I am using is this:\n\n# compile recent git master with threaded delta\ncd git\necho THREADED_DELTA_SEARCH = 1 >>config.mak\nmake install\n\n# get the gcc pack\nmkdir gcc && cd gcc\ngit --bare init\ngit config remote.gcc.url git://git.infradead.org/gcc.git\ngit config remote.gcc.fetch \\\n  '+refs/remotes/gcc.gnu.org/*:refs/remotes/gcc.gnu.org/*'\ngit remote update\n\n# make a copy, so we can run further tests from a known point\ncd ..\ncp -a gcc test\n\n# and test multithreaded large depth/window repacking\ncd test\ngit config pack.threads 4\ntime git repack -a -d -f --window=250 --depth=250\n\n-Peff\n"},{"id":"62247","messageId":"20071207065047.GB13101@coredump.intra.peff.net","threadId":"11153","inReplyTo":"alpine.LFD.0.99999.0712061246120.555@xanadu.home","subject":"Re: Git and GCC","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-07T06:50:47Z","receivedAt":"2007-12-07T06:50:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 06, 2007 at 01:02:58PM -0500, Nicolas Pitre wrote:\n\n> > What is really disappointing is that we saved\n> > only about 20% of the time. I didn't sit around watching the stages, but\n> > my guess is that we spent a long time in the single threaded \"writing\n> > objects\" stage with a thrashing delta cache.\n> \n> Maybe you should run the non threaded repack on the same machine to have \n> a good comparison.\n\nSorry, I should have been more clear. By \"saved\" I meant \"we needed N\nminutes of CPU time, but took only M minutes of real time to use it.\"\nIOW, if we assume that the threading had zero overhead and that we were\ncompletely CPU bound, then the task would have taken N minutes of real\ntime. And obviously those assumptions aren't true, but I was attempting\nto say \"it would have been at most N minutes of real time to do it\nsingle-threaded.\"\n\n> And if you have only 2 CPUs, you will have better performances with\n> pack.threads = 2, otherwise there'll be wasteful task switching going\n> on.\n\nYes, but balanced by one thread running out of data way earlier than the\nother, and completing the task with only one CPU. I am doing a 4-thread\ntest on a quad-CPU right now, and I will also try it with threads=1 and\nthreads=6 for comparison.\n\n> And of course, if the delta cache is being trashed, that might be due to \n> the way the existing pack was previously packed.  Hence the current pack \n> might impact object _access_ when repacking them.  So for a really \n> really fair performance comparison, you'd have to preserve the original \n> pack and swap it back before each repack attempt.\n\nI am working each time from the pack generated by fetching from\ngit://git.infradead.org/gcc.git.\n\n-Peff\n"},{"id":"62250","messageId":"9e4733910712062308t22258c6anb685b18a663e0a31@mail.gmail.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712062120100.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-07T07:08:09Z","receivedAt":"2007-12-07T07:08:09Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/7/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Thu, 6 Dec 2007, Jon Smirl wrote:\n> > >\n> > >         time git blame -C gcc/regclass.c > /dev/null\n> >\n> > jonsmirl@terra:/video/gcc$ time git blame -C gcc/regclass.c > /dev/null\n> >\n> > real    1m21.967s\n> > user    1m21.329s\n>\n> Well, I was also hoping for a \"compared to not-so-aggressive packing\"\n> number on the same machine.. IOW, what I was wondering is whether there is\n> a visible performance downside to the deeper delta chains in the 300MB\n> pack vs the (less aggressive) 500MB pack.\n\nSame machine with a default pack\n\njonsmirl@terra:/video/gcc/.git/objects/pack$ ls -l\ntotal 2145716\n-r--r--r-- 1 jonsmirl jonsmirl   23667932 2007-12-07 02:03\npack-bd163555ea9240a7fdd07d2708a293872665f48b.idx\n-r--r--r-- 1 jonsmirl jonsmirl 2171385413 2007-12-07 02:03\npack-bd163555ea9240a7fdd07d2708a293872665f48b.pack\njonsmirl@terra:/video/gcc/.git/objects/pack$\n\nDelta lengths have virtually no impact. The bigger pack file causes\nmore IO which offsets the increased delta processing time.\n\nOne of my rules is smaller is almost always better. Smaller eliminates\nIO and helps with the CPU cache. It's like the kernel being optimized\nfor size instead of speed ending up being  faster.\n\ntime git blame -C gcc/regclass.c > /dev/null\nreal    1m19.289s\nuser    1m17.853s\nsys     0m0.952s\n\n\n\n>\n>                 Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62251","messageId":"9e4733910712062310s30153afibc44a5550fd9ea99@mail.gmail.com","threadId":"11153","inReplyTo":"20071207063848.GA13101@coredump.intra.peff.net","subject":"Re: Git and GCC","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-07T07:10:49Z","receivedAt":"2007-12-07T07:10:49Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/7/07, Jeff King <peff@peff.net> wrote:\n> On Thu, Dec 06, 2007 at 07:31:21PM -0800, David Miller wrote:\n>\n> > > So it is about 5% bigger. What is really disappointing is that we saved\n> > > only about 20% of the time. I didn't sit around watching the stages, but\n> > > my guess is that we spent a long time in the single threaded \"writing\n> > > objects\" stage with a thrashing delta cache.\n> >\n> > If someone can give me a good way to run this test case I can\n> > have my 64-cpu Niagara-2 box crunch on this and see how fast\n> > it goes and how much larger the resulting pack file is.\n>\n> That would be fun to see. The procedure I am using is this:\n>\n> # compile recent git master with threaded delta\n> cd git\n> echo THREADED_DELTA_SEARCH = 1 >>config.mak\n> make install\n>\n> # get the gcc pack\n> mkdir gcc && cd gcc\n> git --bare init\n> git config remote.gcc.url git://git.infradead.org/gcc.git\n> git config remote.gcc.fetch \\\n>   '+refs/remotes/gcc.gnu.org/*:refs/remotes/gcc.gnu.org/*'\n> git remote update\n>\n> # make a copy, so we can run further tests from a known point\n> cd ..\n> cp -a gcc test\n>\n> # and test multithreaded large depth/window repacking\n> cd test\n> git config pack.threads 4\n\n64 threads with 64 CPUs, if they are multicore you want even more.\nyou need to adjust chunk_size as mentioned in the other mail.\n\n\n> time git repack -a -d -f --window=250 --depth=250\n>\n> -Peff\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62253","messageId":"20071207072710.GA13620@coredump.intra.peff.net","threadId":"11153","inReplyTo":"20071207065047.GB13101@coredump.intra.peff.net","subject":"Re: Git and GCC","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-07T07:27:10Z","receivedAt":"2007-12-07T07:27:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Dec 07, 2007 at 01:50:47AM -0500, Jeff King wrote:\n\n> Yes, but balanced by one thread running out of data way earlier than the\n> other, and completing the task with only one CPU. I am doing a 4-thread\n> test on a quad-CPU right now, and I will also try it with threads=1 and\n> threads=6 for comparison.\n\nHmm. As this has been running, I read the rest of the thread, and it\nlooks like Jon Smirl has already posted the interesting numbers. So\nnevermind, unless there is something particular you would like to see.\n\n-Peff\n"},{"id":"62254","messageId":"20071207073109.GA13638@coredump.intra.peff.net","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712061030560.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-12-07T07:31:09Z","receivedAt":"2007-12-07T07:31:09Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Dec 06, 2007 at 10:35:22AM -0800, Linus Torvalds wrote:\n\n> > What is really disappointing is that we saved only about 20% of the \n> > time. I didn't sit around watching the stages, but my guess is that we \n> > spent a long time in the single threaded \"writing objects\" stage with a \n> > thrashing delta cache.\n> \n> I don't think you spent all that much time writing the objects. That part \n> isn't very intensive, it's mostly about the IO.\n\nIt can get nasty with super-long deltas thrashing the cache, I think.\nBut in this case, I think it ended up being just a poor division of\nlabor caused by the chunk_size parameter using the quite large window\nsize (see elsewhere in the thread for discussion).\n\n> I suspect you may simply be dominated by memory-throughput issues. The \n> delta matching doesn't cache all that well, and using two or more cores \n> isn't going to help all that much if they are largely waiting for memory \n> (and quite possibly also perhaps fighting each other for a shared cache? \n> Is this a Core 2 with the shared L2?)\n\nI think the chunk_size more or less explains it. I have had reasonable\nsuccess keeping both CPUs busy on similar tasks in the past (but with\nsmaller window sizes).\n\nFor reference, it was a Core 2 Duo; do they all share L2, or is there\nsomething I can look for in /proc/cpuinfo?\n\n-Peff\n"},{"id":"62282","messageId":"20071207.045329.204650714.davem@davemloft.net","threadId":"11153","inReplyTo":"9e4733910712062310s30153afibc44a5550fd9ea99@mail.gmail.com","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-07T12:53:29Z","receivedAt":"2007-12-07T12:53:29Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: \"Jon Smirl\" <jonsmirl@gmail.com>\nDate: Fri, 7 Dec 2007 02:10:49 -0500\n\n> On 12/7/07, Jeff King <peff@peff.net> wrote:\n> > On Thu, Dec 06, 2007 at 07:31:21PM -0800, David Miller wrote:\n> >\n> > # and test multithreaded large depth/window repacking\n> > cd test\n> > git config pack.threads 4\n> \n> 64 threads with 64 CPUs, if they are multicore you want even more.\n> you need to adjust chunk_size as mentioned in the other mail.\n\nIt's an 8 core system with 64 cpu threads.\n\n> > time git repack -a -d -f --window=250 --depth=250\n\nDidn't work very well, even with the one-liner patch for\nchunk_size it died.  I think I need to build 64-bit\nbinaries.\n\ndavem@huronp11:~/src/GCC/git/test$ time git repack -a -d -f --window=250 --depth=250\nCounting objects: 1190671, done.\nfatal: Out of memory? mmap failed: Cannot allocate memory\n\nreal    58m36.447s\nuser    289m8.270s\nsys     4m40.680s\ndavem@huronp11:~/src/GCC/git/test$ \n\nWhile it did run the load was anywhere between 5 and 9, although it\ndid create 64 threads, and the size of the process was about 3.2GB\nThis may be in part why it wasn't able to use all 64 thread\neffectively.  Like I said it seemed to have 9 active at best, at any\none time, most of the time only 4 or 5 were busy doing anything.\n\nAlso I could end up being performance limited by SHA, it's not very\nwell tuned on Sparc.  It's been on my TODO list to code up the crypto\nunit support for Niagara-2 in the kernel, then work with Herbert Xu on\nthe userland interfaces to take advantage of that in things like\nlibssl.  Even a better C/asm version would probably improve GIT\nperformance a bit.\n\nIs SHA a significant portion of the compute during these repacks?\nI should run oprofile...\n"},{"id":"62298","messageId":"alpine.LFD.0.9999.0712070919590.7274@woody.linux-foundation.org","threadId":"11153","inReplyTo":"20071207.045329.204650714.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-07T17:23:47Z","receivedAt":"2007-12-07T17:23:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 7 Dec 2007, David Miller wrote:\n> \n> Also I could end up being performance limited by SHA, it's not very\n> well tuned on Sparc.  It's been on my TODO list to code up the crypto\n> unit support for Niagara-2 in the kernel, then work with Herbert Xu on\n> the userland interfaces to take advantage of that in things like\n> libssl.  Even a better C/asm version would probably improve GIT\n> performance a bit.\n\nI doubt yu can use the hardware support. Kernel-only hw support is \ninherently broken for any sane user-space usage, the setup costs are just \nway way too high. To be useful, crypto engines need to support direct user \nspace access (ie a regular instruction, with all state being held in \nnormal registers that get saved/restored by the kernel).\n\n> Is SHA a significant portion of the compute during these repacks?\n> I should run oprofile...\n\nSHA1 is almost totally insignificant on x86. It hardly shows up. But we \nhave a good optimized version there.\n\nzlib tends to be a lot more noticeable (especially the uncompression: it \nmay be faster than compression, but it's done _so_ much more that it \ntotally dominates).\n\n\t\t\tLinus\n"},{"id":"62308","messageId":"alpine.LFD.0.99999.0712071419390.555@xanadu.home","threadId":"11153","inReplyTo":"9e4733910712062308t22258c6anb685b18a663e0a31@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-07T19:36:12Z","receivedAt":"2007-12-07T19:36:12Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Dec 2007, Jon Smirl wrote:\n\n> On 12/7/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> >\n> >\n> > On Thu, 6 Dec 2007, Jon Smirl wrote:\n> > > >\n> > > >         time git blame -C gcc/regclass.c > /dev/null\n> > >\n> > > jonsmirl@terra:/video/gcc$ time git blame -C gcc/regclass.c > /dev/null\n> > >\n> > > real    1m21.967s\n> > > user    1m21.329s\n> >\n> > Well, I was also hoping for a \"compared to not-so-aggressive packing\"\n> > number on the same machine.. IOW, what I was wondering is whether there is\n> > a visible performance downside to the deeper delta chains in the 300MB\n> > pack vs the (less aggressive) 500MB pack.\n> \n> Same machine with a default pack\n> \n> jonsmirl@terra:/video/gcc/.git/objects/pack$ ls -l\n> total 2145716\n> -r--r--r-- 1 jonsmirl jonsmirl   23667932 2007-12-07 02:03\n> pack-bd163555ea9240a7fdd07d2708a293872665f48b.idx\n> -r--r--r-- 1 jonsmirl jonsmirl 2171385413 2007-12-07 02:03\n> pack-bd163555ea9240a7fdd07d2708a293872665f48b.pack\n> jonsmirl@terra:/video/gcc/.git/objects/pack$\n> \n> Delta lengths have virtually no impact. \n\nI can confirm this.\n\nI just did a repack keeping the default depth of 50 but with window=100 \ninstead of the default of 10, and the pack shrunk from 2171385413 bytes \ndown to 410607140 bytes.\n\nSo our default window size is definitely not adequate for the gcc repo.\n\nOTOH, I recall tytso mentioning something about not having much return \non  a bigger window size in his tests when he proposed to increase the \ndefault delta depth to 50.  So there is definitely some kind of threshold \nat which point the increased window size stops being advantageous wrt \nthe number of cycles involved, and we should find a way to correlate it \nto the data set to have a better default window size than the current \nfixed default.\n\n\nNicolas\n"},{"id":"62314","messageId":"4759AC8E.3070102@develer.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712070919590.7274@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Giovanni Bajo","fromEmail":"rasky@develer.com","sentAt":"2007-12-07T20:26:54Z","receivedAt":"2007-12-07T20:26:54Z","isPatch":false,"sender":{"key":"rasky@develer.com","avatar":"https://gravatar.com/avatar/852c310d41ec49adc82a6e2937c16aa0dc87bad4da1be1d686da48aa255b875e?d=mp&s=160"},"body":"On 12/7/2007 6:23 PM, Linus Torvalds wrote:\n\n>> Is SHA a significant portion of the compute during these repacks?\n>> I should run oprofile...\n> \n> SHA1 is almost totally insignificant on x86. It hardly shows up. But we \n> have a good optimized version there.\n> \n> zlib tends to be a lot more noticeable (especially the uncompression: it \n> may be faster than compression, but it's done _so_ much more that it \n> totally dominates).\n\nHave you considered alternatives, like:\nhttp://www.oberhumer.com/opensource/ucl/\n-- \nGiovanni Bajo\n"},{"id":"62343","messageId":"m3hciutaoq.fsf@roke.D-201","threadId":"11153","inReplyTo":"4759AC8E.3070102@develer.com","subject":"Re: Git and GCC","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-07T22:14:15Z","receivedAt":"2007-12-07T22:14:15Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Giovanni Bajo <rasky@develer.com> writes:\n\n> On 12/7/2007 6:23 PM, Linus Torvalds wrote:\n> \n> >> Is SHA a significant portion of the compute during these repacks?\n> >> I should run oprofile...\n> > SHA1 is almost totally insignificant on x86. It hardly shows up. But\n> > we have a good optimized version there.\n> > zlib tends to be a lot more noticeable (especially the\n> > *uncompression*: it may be faster than compression, but it's done _so_\n> > much more that it totally dominates).\n> \n> Have you considered alternatives, like:\n> http://www.oberhumer.com/opensource/ucl/\n\n<quote>\n  As compared to LZO, the UCL algorithms achieve a better compression\n  ratio but *decompression* is a little bit slower. See below for some\n  rough timings.\n</quote>\n\nIt is uncompression speed that is more important, because it is used\nmuch more often.\n\n-- \nJakub Narebski\nShadeHawk on #git\n"},{"id":"62345","messageId":"B2B3D9E9-4A5C-4DE0-9FF9-B0F9304BDCD2@vicaya.com","threadId":"11153","inReplyTo":"m3hciutaoq.fsf@roke.D-201","subject":"Re: Git and GCC","fromName":"Luke Lu","fromEmail":"git@vicaya.com","sentAt":"2007-12-07T23:04:25Z","receivedAt":"2007-12-07T23:04:25Z","isPatch":false,"sender":{"key":"git@vicaya.com","avatar":null},"body":"On Dec 7, 2007, at 2:14 PM, Jakub Narebski wrote:\n> Giovanni Bajo <rasky@develer.com> writes:\n>> On 12/7/2007 6:23 PM, Linus Torvalds wrote:\n>>>> Is SHA a significant portion of the compute during these repacks?\n>>>> I should run oprofile...\n>>> SHA1 is almost totally insignificant on x86. It hardly shows up. But\n>>> we have a good optimized version there.\n>>> zlib tends to be a lot more noticeable (especially the\n>>> *uncompression*: it may be faster than compression, but it's done  \n>>> _so_\n>>> much more that it totally dominates).\n>>\n>> Have you considered alternatives, like:\n>> http://www.oberhumer.com/opensource/ucl/\n>\n> <quote>\n>   As compared to LZO, the UCL algorithms achieve a better compression\n>   ratio but *decompression* is a little bit slower. See below for some\n>   rough timings.\n> </quote>\n>\n> It is uncompression speed that is more important, because it is used\n> much more often.\n\nSo why didn't we consider lzo then? It's much faster than zlib.\n\n__Luke\n\n  \n"},{"id":"62347","messageId":"1197069298.6118.1.camel@ozzu","threadId":"11153","inReplyTo":"m3hciutaoq.fsf@roke.D-201","subject":"Re: Git and GCC","fromName":"Giovanni Bajo","fromEmail":"rasky@develer.com","sentAt":"2007-12-07T23:14:58Z","receivedAt":"2007-12-07T23:14:58Z","isPatch":false,"sender":{"key":"rasky@develer.com","avatar":"https://gravatar.com/avatar/852c310d41ec49adc82a6e2937c16aa0dc87bad4da1be1d686da48aa255b875e?d=mp&s=160"},"body":"On Fri, 2007-12-07 at 14:14 -0800, Jakub Narebski wrote:\n\n> > >> Is SHA a significant portion of the compute during these repacks?\n> > >> I should run oprofile...\n> > > SHA1 is almost totally insignificant on x86. It hardly shows up. But\n> > > we have a good optimized version there.\n> > > zlib tends to be a lot more noticeable (especially the\n> > > *uncompression*: it may be faster than compression, but it's done _so_\n> > > much more that it totally dominates).\n> > \n> > Have you considered alternatives, like:\n> > http://www.oberhumer.com/opensource/ucl/\n> \n> <quote>\n>   As compared to LZO, the UCL algorithms achieve a better compression\n>   ratio but *decompression* is a little bit slower. See below for some\n>   rough timings.\n> </quote>\n> \n> It is uncompression speed that is more important, because it is used\n> much more often.\n\nI know, but the point is not what is the fastestest, but if it's fast\nenough to get off the profiles. I think UCL is fast enough since it's\nstill times faster than zlib. Anyway, LZO is GPL too, so why not\nconsidering it too. They are good libraries.\n-- \nGiovanni Bajo\n"},{"id":"62348","messageId":"4aca3dc20712071533k3189d25dp901c5941e5326ead@mail.gmail.com","threadId":"11153","inReplyTo":"1197069298.6118.1.camel@ozzu","subject":"Re: Git and GCC","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-07T23:33:26Z","receivedAt":"2007-12-07T23:33:26Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/7/07, Giovanni Bajo <rasky@develer.com> wrote:\n> On Fri, 2007-12-07 at 14:14 -0800, Jakub Narebski wrote:\n>\n> > > >> Is SHA a significant portion of the compute during these repacks?\n> > > >> I should run oprofile...\n> > > > SHA1 is almost totally insignificant on x86. It hardly shows up. But\n> > > > we have a good optimized version there.\n> > > > zlib tends to be a lot more noticeable (especially the\n> > > > *uncompression*: it may be faster than compression, but it's done _so_\n> > > > much more that it totally dominates).\n> > >\n> > > Have you considered alternatives, like:\n> > > http://www.oberhumer.com/opensource/ucl/\n> >\n> > <quote>\n> >   As compared to LZO, the UCL algorithms achieve a better compression\n> >   ratio but *decompression* is a little bit slower. See below for some\n> >   rough timings.\n> > </quote>\n> >\n> > It is uncompression speed that is more important, because it is used\n> > much more often.\n>\n> I know, but the point is not what is the fastestest, but if it's fast\n> enough to get off the profiles. I think UCL is fast enough since it's\n> still times faster than zlib. Anyway, LZO is GPL too, so why not\n> considering it too. They are good libraries.\n\n\nAt worst, you could also use fastlz (www.fastlz.org), which is faster\nthan all of these by a factor of 4 (and compression wise, is actually\nsometimes better, sometimes worse, than LZO).\n"},{"id":"62350","messageId":"1197074839.22471.34.camel@brick","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712061030560.13796@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-08T00:47:19Z","receivedAt":"2007-12-08T00:47:19Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"Some interesting stats from the highly packed gcc repo.  The long chain\nlengths very quickly tail off.  Over 60% of the objects have a chain\nlength of 20 or less.  If anyone wants the full list let me know.  I\nalso have included a few other interesting points, the git default\ndepth of 50, my initial guess of 100 and every 10% in the cumulative\ndistribution from 60-100%.\n\nThis shows the git default of 50 really isn't that bad, and after\nabout 100 it really starts to get sparse.  \n\nHarvey\n\n1:\t103817\t103817\t10.20%\t1017922\n2:\t67332\t171149\t16.81%\n3:\t57520\t228669\t22.46%\n4:\t52570\t281239\t27.63%\n5:\t43910\t325149\t31.94%\n6:\t37520\t362669\t35.63%\n7:\t35248\t397917\t39.09%\n8:\t29819\t427736\t42.02%\n9:\t27619\t455355\t44.73%\n10:\t22656\t478011\t46.96%\n11:\t21073\t499084\t49.03%\n12:\t18738\t517822\t50.87%\n13:\t16674\t534496\t52.51%\n14:\t14882\t549378\t53.97%\n15:\t14424\t563802\t55.39%\n16:\t12765\t576567\t56.64%\n17:\t11662\t588229\t57.79%\n18:\t11845\t600074\t58.95%\n19:\t11694\t611768\t60.10%\n20:\t9625\t621393\t61.05%\n34:\t5354\t719356\t70.67%\n50:\t3395\t785342\t77.15%\n60:\t2547\t815072\t80.07%\n100:\t1644\t898284\t88.25%\n113:\t1292\t917046\t90.09%\n158:\t959\t967429\t95.04%\n200:\t652\t997653\t98.01%\n219:\t491\t1008132\t99.04%\n245:\t179\t1017717\t99.98%\n246:\t111\t1017828\t99.99%\n247:\t61\t1017889\t100.00%\n248:\t27\t1017916\t100.00%\n249:\t6\t1017922\t100.00%\n"},{"id":"62356","messageId":"20071207.175529.104353710.davem@davemloft.net","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712070919590.7274@woody.linux-foundation.org","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-08T01:55:29Z","receivedAt":"2007-12-08T01:55:29Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: Linus Torvalds <torvalds@linux-foundation.org>\nDate: Fri, 7 Dec 2007 09:23:47 -0800 (PST)\n\n> \n> \n> On Fri, 7 Dec 2007, David Miller wrote:\n> > \n> > Also I could end up being performance limited by SHA, it's not very\n> > well tuned on Sparc.  It's been on my TODO list to code up the crypto\n> > unit support for Niagara-2 in the kernel, then work with Herbert Xu on\n> > the userland interfaces to take advantage of that in things like\n> > libssl.  Even a better C/asm version would probably improve GIT\n> > performance a bit.\n> \n> I doubt yu can use the hardware support. Kernel-only hw support is \n> inherently broken for any sane user-space usage, the setup costs are just \n> way way too high. To be useful, crypto engines need to support direct user \n> space access (ie a regular instruction, with all state being held in \n> normal registers that get saved/restored by the kernel).\n\nUnfortunately they are hypervisor calls, and you have to give\nthe thing physical addresses for the buffer to work on, so\nletting userland get at it directly isn't currently doable.\n\nI still believe that there are cases where userland can take\nadvantage of in-kernel crypto devices, such as when we are\nstreaming the data into the kernel anyways (for a write()\nor sendmsg()) and the user just wants the transformation to\nbe done on that stream.\n\nAs a specific case, hardware crypto SSL support works quite\nwell for sendmsg() user packet data.  And this the kind of API\nSolaris provides to get good SSL performance with Niagara.\n\n> > Is SHA a significant portion of the compute during these repacks?\n> > I should run oprofile...\n> \n> SHA1 is almost totally insignificant on x86. It hardly shows up. But we \n> have a good optimized version there.\n\nOk.\n\n> zlib tends to be a lot more noticeable (especially the uncompression: it \n> may be faster than compression, but it's done _so_ much more that it \n> totally dominates).\n\nzlib is really hard to optimize on Sparc, I've tried numerous times.\nActually compress is the real cycle killer, and in that case the inner\nloop wants to dereference 2-byte shorts at a time but they are\nunaligned half of the time, and any the check for alignment nullifies\nthe gains of avoiding the two byte loads.\n\nUncompress I don't think is optimized at all on any platform with\nasm stuff like the compress side is.  It's a pretty straightforward\ntransformation and the memory accesses dominate the overhead.\n\nI'll do some profiling to see what might be worth looking into.\n"},{"id":"62409","messageId":"Pine.LNX.4.64.0712081156450.27959@racer.site","threadId":"11153","inReplyTo":"4aca3dc20712071533k3189d25dp901c5941e5326ead@mail.gmail.com","subject":"Re: Git and GCC","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-12-08T12:00:41Z","receivedAt":"2007-12-08T12:00:41Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 7 Dec 2007, Daniel Berlin wrote:\n\n> On 12/7/07, Giovanni Bajo <rasky@develer.com> wrote:\n> > On Fri, 2007-12-07 at 14:14 -0800, Jakub Narebski wrote:\n> >\n> > > > >> Is SHA a significant portion of the compute during these \n> > > > >> repacks? I should run oprofile...\n> > > > > SHA1 is almost totally insignificant on x86. It hardly shows up. \n> > > > > But we have a good optimized version there. zlib tends to be a \n> > > > > lot more noticeable (especially the *uncompression*: it may be \n> > > > > faster than compression, but it's done _so_ much more that it \n> > > > > totally dominates).\n> > > >\n> > > > Have you considered alternatives, like: \n> > > > http://www.oberhumer.com/opensource/ucl/\n> > >\n> > > <quote>\n> > >   As compared to LZO, the UCL algorithms achieve a better \n> > >   compression ratio but *decompression* is a little bit slower. See \n> > >   below for some rough timings.\n> > > </quote>\n> > >\n> > > It is uncompression speed that is more important, because it is used \n> > > much more often.\n> >\n> > I know, but the point is not what is the fastestest, but if it's fast \n> > enough to get off the profiles. I think UCL is fast enough since it's \n> > still times faster than zlib. Anyway, LZO is GPL too, so why not \n> > considering it too. They are good libraries.\n> \n> \n> At worst, you could also use fastlz (www.fastlz.org), which is faster \n> than all of these by a factor of 4 (and compression wise, is actually \n> sometimes better, sometimes worse, than LZO).\n\nfastLZ is awfully short on details when it comes to a comparison of the \nresulting file sizes.\n\nThe only result I saw was that for the (single) example they chose, \ncompressed size was 470MB as opposed to 361MB for zip's _fastest_ mode.\n\nReally, that's not acceptable for me in the context of git.\n\nBesides, if you change the compression algorithm you will have to add \nsupport for legacy clients to _recompress_ with libz.  Which most likely \nwould make Sisyphos grin watching them servers.\n\nCiao,\nDscho\n"},{"id":"62556","messageId":"20071210095426.GA32611@iram.es","threadId":"11153","inReplyTo":"1197074839.22471.34.camel@brick","subject":"Re: Git and GCC","fromName":"Gabriel Paubert","fromEmail":"paubert@iram.es","sentAt":"2007-12-10T09:54:26Z","receivedAt":"2007-12-10T09:54:26Z","isPatch":false,"sender":{"key":"paubert@iram.es","avatar":null},"body":"On Fri, Dec 07, 2007 at 04:47:19PM -0800, Harvey Harrison wrote:\n> Some interesting stats from the highly packed gcc repo.  The long chain\n> lengths very quickly tail off.  Over 60% of the objects have a chain\n> length of 20 or less.  If anyone wants the full list let me know.  I\n> also have included a few other interesting points, the git default\n> depth of 50, my initial guess of 100 and every 10% in the cumulative\n> distribution from 60-100%.\n> \n> This shows the git default of 50 really isn't that bad, and after\n> about 100 it really starts to get sparse.  \n\nDo you have a way to know which files have the longest chains?\n\nI have a suspiscion that the ChangeLog* files are among them,\nnot only because they are, almost without exception, only modified\nby prepending text to the previous version (and a fairly small amount\ncompared to the size of the file), and therefore the diff is simple\n(a single hunk) so that the limit on chain depth is probably what\ncauses a new copy to be created. \n\nBesides that these files grow quite large and become some of the \nlargest files in the tree, and at least one of them is changed \nfor every commit. This leads again to many versions of fairly \nlarge files.\n\nIf this guess is right, this implies that most of the size gains\nfrom longer chains comes from having less copies of the ChangeLog*\nfiles. From a performance point of view, it is rather favourable\nsince the differences are simple. This would also explain why\nthe window parameter has little effect.\n\n\tRegards,\n\tGabriel\n"},{"id":"62558","messageId":"20071210.015749.204978503.davem@davemloft.net","threadId":"11153","inReplyTo":"20071207.045329.204650714.davem@davemloft.net","subject":"Re: Git and GCC","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-10T09:57:49Z","receivedAt":"2007-12-10T09:57:49Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: David Miller <davem@davemloft.net>\nDate: Fri, 07 Dec 2007 04:53:29 -0800 (PST)\n\n> I should run oprofile...\n\nWhile doing the initial object counting, most of the time is spent in\nlookup_object(), memcmp() (via hashcmp()), and inflate().  I tried to\nsee if I could do some tricks on sparc with the hashcmp() but the sha1\npointers are very often not even 4 byte aligned.\n\nI suspect lookup_object() could be improved if it didn't use a hash\ntable without chaining, but I can see why 'struct object' size is a\nconcern and thus why things are done the way they are.\n\nsamples  %        app name                 symbol name\n504      13.7517  libc-2.6.1.so            memcmp\n386      10.5321  libz.so.1.2.3.3          inflate\n288       7.8581  git                      lookup_object\n248       6.7667  libz.so.1.2.3.3          inflate_fast\n201       5.4843  libz.so.1.2.3.3          inflate_table\n175       4.7749  git                      decode_tree_entry\n ...\n\nDeltifying is %94 consumed by create_delta(), the rest is completely\nin the noise.\n\nsamples  %        app name                 symbol name\n10581    94.8373  git                      create_delta\n181       1.6223  git                      create_delta_index\n72        0.6453  git                      prepare_pack\n55        0.4930  libc-2.6.1.so            loop\n34        0.3047  libz.so.1.2.3.3          inflate_fast\n33        0.2958  libc-2.6.1.so            _int_malloc\n22        0.1972  libshadow.so             shadowUpdatePacked\n21        0.1882  libc-2.6.1.so            _int_free\n19        0.1703  libc-2.6.1.so            malloc\n ...\n"},{"id":"62583","messageId":"alpine.LFD.0.99999.0712101022300.555@xanadu.home","threadId":"11153","inReplyTo":"20071210095426.GA32611@iram.es","subject":"Re: Git and GCC","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-10T15:35:40Z","receivedAt":"2007-12-10T15:35:40Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 10 Dec 2007, Gabriel Paubert wrote:\n\n> On Fri, Dec 07, 2007 at 04:47:19PM -0800, Harvey Harrison wrote:\n> > Some interesting stats from the highly packed gcc repo.  The long chain\n> > lengths very quickly tail off.  Over 60% of the objects have a chain\n> > length of 20 or less.  If anyone wants the full list let me know.  I\n> > also have included a few other interesting points, the git default\n> > depth of 50, my initial guess of 100 and every 10% in the cumulative\n> > distribution from 60-100%.\n> > \n> > This shows the git default of 50 really isn't that bad, and after\n> > about 100 it really starts to get sparse.  \n> \n> Do you have a way to know which files have the longest chains?\n\nWith 'git verify-pack -v' you get the delta depth for each object.\nThen you can use 'git show' with the object SHA1 to see its content.\n\n> I have a suspiscion that the ChangeLog* files are among them,\n> not only because they are, almost without exception, only modified\n> by prepending text to the previous version (and a fairly small amount\n> compared to the size of the file), and therefore the diff is simple\n> (a single hunk) so that the limit on chain depth is probably what\n> causes a new copy to be created. \n\nMy gcc repo is currently repacked with a max delta depth of 50, and \na quick sample of those objects at the depth limit does indeed show the \ncontent of the ChangeLog file.  But I have occurrences of the root \ndirectory tree object too, and the \"GCC machine description for IA-32\" \ncontent as well.\n\nBut yes, the really deep delta chains are most certainly going to \ncontain those ChangeLog files.\n\n> Besides that these files grow quite large and become some of the \n> largest files in the tree, and at least one of them is changed \n> for every commit. This leads again to many versions of fairly \n> large files.\n> \n> If this guess is right, this implies that most of the size gains\n> from longer chains comes from having less copies of the ChangeLog*\n> files. From a performance point of view, it is rather favourable\n> since the differences are simple. This would also explain why\n> the window parameter has little effect.\n\nWell, actually the window parameter does have big effects.  For instance \nthe default of 10 is completely inadequate for the gcc repo, since \nchanging the window size from 10 to 100 made the corresponding pack \nshrink from 2.1GB down to 400MB, with the same max delta depth.\n\n\nNicolas\n"},{"id":"63495","messageId":"37BDCA73-4318-4BC8-9CCE-1DA30E4A09FC@adacore.com","threadId":"11153","inReplyTo":"1197572755.898.15.camel@brick","subject":"\"Argument list too long\" in git remote update (Was: Git and GCC)","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2007-12-17T22:15:06Z","receivedAt":"2007-12-17T22:15:06Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\t\t\nOn Dec 13, 2007, at 14:05, Harvey Harrison wrote:\n> After the discussions lately regarding the gcc svn mirror.  I'm coming\n> up with a recipe to set up your own git-svn mirror.  Suggestions on  \n> the\n> following.\n>\n> // Create directory and initialize git\n> mkdir gcc\n> cd gcc\n> git init\n> // add the remote site that currently mirrors gcc\n> // I have chosen the name gcc.gnu.org *1* as my local name to refer to\n> // this choose something else if you like\n> git remote add gcc.gnu.org git://git.infradead.org/gcc.git\n> // fetching someone else's remote branches is not a standard thing  \n> to do\n> // so we'll need to edit our .git/config file\n> // you should have a section that looks like:\n> [remote \"gcc.gnu.org\"]\n>        url = git://git.infradead.org/gcc.git\n>        fetch = +refs/heads/*:refs/remotes/gcc.gnu.org/*\n> // infradead's mirror puts the gcc svn branches in its own namespace\n> // refs/remotes/gcc.gnu.org/*\n> // change our fetch line accordingly\n> [remote \"gcc.gnu.org\"]\n>        url = git://git.infradead.org/gcc.git\n>        fetch = +refs/remotes/gcc.gnu.org/*:refs/remotes/gcc.gnu.org/*\n> // fetch the remote data from the mirror site\n> git remote update\n\nWith git version 1.5.3.6 on Mac OS X, this results in:\npotomac%:~/gcc%git remote update\nUpdating gcc.gnu.org\n/opt/git/bin/git-fetch: line 220: /opt/git/bin/git: Argument list too  \nlong\nwarning: no common commits\n[after a long wait and a good amount of network traffic]\nfatal: index-pack died of signal 13\nfetch gcc.gnu.org: command returned error: 126\npotomac%:~/gcc%\n\nAny ideas on what to do to resolve this?\n"},{"id":"63508","messageId":"Pine.LNX.4.64.0712172257380.9446@racer.site","threadId":"11153","inReplyTo":"37BDCA73-4318-4BC8-9CCE-1DA30E4A09FC@adacore.com","subject":"Re: \"Argument list too long\" in git remote update (Was: Git and GCC)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-12-17T22:59:57Z","receivedAt":"2007-12-17T22:59:57Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 17 Dec 2007, Geert Bosch wrote:\n\n> \t\tOn Dec 13, 2007, at 14:05, Harvey Harrison wrote:\n> > After the discussions lately regarding the gcc svn mirror.  I'm coming\n> > up with a recipe to set up your own git-svn mirror.  Suggestions on the\n> > following.\n> > \n> > // Create directory and initialize git\n> > mkdir gcc\n> > cd gcc\n> > git init\n> > // add the remote site that currently mirrors gcc\n> > // I have chosen the name gcc.gnu.org *1* as my local name to refer to\n> > // this choose something else if you like\n> > git remote add gcc.gnu.org git://git.infradead.org/gcc.git\n> > // fetching someone else's remote branches is not a standard thing to do\n> > // so we'll need to edit our .git/config file\n> > // you should have a section that looks like:\n> > [remote \"gcc.gnu.org\"]\n> >       url = git://git.infradead.org/gcc.git\n> >       fetch = +refs/heads/*:refs/remotes/gcc.gnu.org/*\n> > // infradead's mirror puts the gcc svn branches in its own namespace\n> > // refs/remotes/gcc.gnu.org/*\n> > // change our fetch line accordingly\n> > [remote \"gcc.gnu.org\"]\n> >       url = git://git.infradead.org/gcc.git\n> >       fetch = +refs/remotes/gcc.gnu.org/*:refs/remotes/gcc.gnu.org/*\n> > // fetch the remote data from the mirror site\n> > git remote update\n> \n> With git version 1.5.3.6 on Mac OS X, this results in:\n> potomac%:~/gcc%git remote update\n> Updating gcc.gnu.org\n> /opt/git/bin/git-fetch: line 220: /opt/git/bin/git: Argument list too long\n> warning: no common commits\n> [after a long wait and a good amount of network traffic]\n> fatal: index-pack died of signal 13\n> fetch gcc.gnu.org: command returned error: 126\n> potomac%:~/gcc%\n> \n> Any ideas on what to do to resolve this?\n\nUnfortunately, the builtin remote did not make it into git's master yet, \nand it will probably miss 1.5.4.\n\nChances are that this would make the bug go away, but Junio said that on \none of his machines, the regression tests fail with the builtin remote.\n\nIn the meantime, \"git fetch gcc.gnu.org\" should do what you want, \nmethinks.\n\nHth,\nDscho\n"},{"id":"63510","messageId":"alpine.LFD.0.9999.0712171455220.21557@woody.linux-foundation.org","threadId":"11153","inReplyTo":"37BDCA73-4318-4BC8-9CCE-1DA30E4A09FC@adacore.com","subject":"Re: \"Argument list too long\" in git remote update (Was: Git and GCC)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-17T23:01:25Z","receivedAt":"2007-12-17T23:01:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 17 Dec 2007, Geert Bosch wrote:\n> \n> With git version 1.5.3.6 on Mac OS X, this results in:\n> potomac%:~/gcc%git remote update\n> Updating gcc.gnu.org\n> /opt/git/bin/git-fetch: line 220: /opt/git/bin/git: Argument list too long\n\nOops.\n\n> Any ideas on what to do to resolve this?\n\nCan you try the current git tree? \"git fetch\" is built-in these days, and \nthat old shell-script that ran \"git fetch--tool\" on all the refs is no \nmore, so most likely the problem simply no longer exists.\n\nBut maybe there is some way to raise the argument size limit on OS X. One \nthing to check is whether maybe you have an excessively big environment \n(just run \"printenv\" to see what it contains) which might be cutting down \non the size allowed for arguments.\n\n\t\t\tLinus\n"},{"id":"63545","messageId":"20071218013401.A1716@edi-view2.cisco.com","threadId":"11153","inReplyTo":"alpine.LFD.0.9999.0712171455220.21557@woody.linux-foundation.org","subject":"Re: \"Argument list too long\" in git remote update (Was: Git and GCC)","fromName":"Derek Fawcus","fromEmail":"dfawcus@cisco.com","sentAt":"2007-12-18T01:34:01Z","receivedAt":"2007-12-18T01:34:01Z","isPatch":false,"sender":{"key":"dfawcus@cisco.com","avatar":null},"body":"On Mon, Dec 17, 2007 at 03:01:25PM -0800, Linus Torvalds wrote:\n> \n> > With git version 1.5.3.6 on Mac OS X, this results in:\n> > potomac%:~/gcc%git remote update\n> > Updating gcc.gnu.org\n> > /opt/git/bin/git-fetch: line 220: /opt/git/bin/git: Argument list too long\n\n> But maybe there is some way to raise the argument size limit on OS X.\n\nWell the certification for Leopard claims it can be up to 256k.\n\nI don't know about Tiger or earlier,  but ARG_MAX on my 10.4\nbox is also (256 * 1024).\n\nSo - how much do people want?  Or maybe there is some sort limit in play here?\n\nDF\n"},{"id":"63548","messageId":"20071218015225.GV14735@spearce.org","threadId":"11153","inReplyTo":"20071218013401.A1716@edi-view2.cisco.com","subject":"Re: \"Argument list too long\" in git remote update (Was: Git and GCC)","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-12-18T01:52:25Z","receivedAt":"2007-12-18T01:52:25Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Derek Fawcus <dfawcus@cisco.com> wrote:\n> On Mon, Dec 17, 2007 at 03:01:25PM -0800, Linus Torvalds wrote:\n> > \n> > > With git version 1.5.3.6 on Mac OS X, this results in:\n> > > potomac%:~/gcc%git remote update\n> > > Updating gcc.gnu.org\n> > > /opt/git/bin/git-fetch: line 220: /opt/git/bin/git: Argument list too long\n> \n> > But maybe there is some way to raise the argument size limit on OS X.\n> \n> Well the certification for Leopard claims it can be up to 256k.\n> \n> I don't know about Tiger or earlier,  but ARG_MAX on my 10.4\n> box is also (256 * 1024).\n> \n> So - how much do people want?  Or maybe there is some sort limit in play here?\n\nI heard there's like 1000 branches in that GCC repository, not\ncounting tags.\n\nEach branch name is at least 12 bytes or so (\"refs/heads/..\").\nIt adds up.\n\nIt should be fixed in latest git (1.5.4-rc0) as that uses the new\nbuiltin-fetch implementation, which passes the lists around in\nmemory rather than on the command line.  Much faster, and doesn't\nsuffer from argument/environment limits.\n\n-- \nShawn.\n"},{"id":"108379","messageId":"alpine.DEB.1.00.0903181657180.10279@pacific.mpi-cbg.de","threadId":"11153","inReplyTo":"Pine.LNX.4.64.0712061201580.27959@racer.site","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-03-18T16:01:31Z","receivedAt":"2009-03-18T16:01:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Dec 2007, Johannes Schindelin wrote:\n\n> \n> The default was not to change the window or depth at all.  As suggested\n> by Jon Smirl, Linus Torvalds and others, default to\n> \n> \t--window=250 --depth=250\n> \n> Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> ---\n\nGuess what.  This is still unresolved, and yet somebody else had to be \nbitten by 'git gc --aggressive' being everything but aggressive.\n\nSo...  I think it is high time to resolve the issue, either by applying \nthis patch with a delay of over one year, or by the pack wizards trying to \nimplement that 'never fall back to a worse delta' idea mentioned in this \nthread.\n\nAlthough I suggest, really, that implying --depth=250 --window=250 (unless \noverridden by the config) with --aggressive is not at all wrong.\n\nCiao,\nDscho\n"},{"id":"108383","messageId":"87ocvyvlsd.fsf@iki.fi","threadId":"11153","inReplyTo":"alpine.DEB.1.00.0903181657180.10279@pacific.mpi-cbg.de","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Teemu Likonen","fromEmail":"tlikonen@iki.fi","sentAt":"2009-03-18T16:27:30Z","receivedAt":"2009-03-18T16:27:30Z","isPatch":true,"sender":{"key":"tlikonen@iki.fi","avatar":null},"body":"On 2009-03-18 17:01 (+0100), Johannes Schindelin wrote:\n\n>> The default was not to change the window or depth at all. As\n>> suggested by Jon Smirl, Linus Torvalds and others, default to\n>> \n>> \t--window=250 --depth=250\n\n> Guess what. This is still unresolved, and yet somebody else had to be\n> bitten by 'git gc --aggressive' being everything but aggressive.\n\nPieter de Bie's tests seem to suggest that usually --window=50\n--depth=50 gives about the same results than with higher values:\n\n    http://vcscompare.blogspot.com/2008/06/git-repack-parameters.html\n\nI don't understand the issue very well myself so I really can't say what\nwould be a/the good value. Anyway, I agree that it would be nice if \"git\ngc --aggressive\" were aggressive and a user wouldn't need to know about\n\"git repack\" and its cryptical low-levelish options.\n"},{"id":"108401","messageId":"alpine.LFD.2.00.0903181401000.30483@xanadu.home","threadId":"11153","inReplyTo":"alpine.DEB.1.00.0903181657180.10279@pacific.mpi-cbg.de","subject":"Re: [PATCH] gc --aggressive: make it really aggressive","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-03-18T18:02:39Z","receivedAt":"2009-03-18T18:02:39Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 18 Mar 2009, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Thu, 6 Dec 2007, Johannes Schindelin wrote:\n> \n> > \n> > The default was not to change the window or depth at all.  As suggested\n> > by Jon Smirl, Linus Torvalds and others, default to\n> > \n> > \t--window=250 --depth=250\n> > \n> > Signed-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n> > ---\n> \n> Guess what.  This is still unresolved, and yet somebody else had to be \n> bitten by 'git gc --aggressive' being everything but aggressive.\n> \n> So...  I think it is high time to resolve the issue, either by applying \n> this patch with a delay of over one year, or by the pack wizards trying to \n> implement that 'never fall back to a worse delta' idea mentioned in this \n> thread.\n\nThis is just a bit complicated to implement (cycle avoidance, etc).\n\n> Although I suggest, really, that implying --depth=250 --window=250 (unless \n> overridden by the config) with --aggressive is not at all wrong.\n\nACK.\n\n\nNicolas\n"}]}