{"thread":{"id":"7934","subject":"Git benchmarks at OpenOffice.org wiki","startedAt":"2007-05-01T21:46:14Z","lastAt":"2007-05-07T15:22:52Z","messageCount":36,"participants":["Jakub Narebski","Junio C Hamano","Andy Parkins","Julian Phillips","Johannes Schindelin","Jan Holesovsky","Petr Baudis","Florian Weimer","Robin Rosenberg","Martin Langhoff","Alex Riesen","Johannes Sixt","Linus Torvalds"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"40860","messageId":"200705012346.14997.jnareb@gmail.com","threadId":"7934","inReplyTo":null,"subject":"Git benchmarks at OpenOffice.org wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-01T21:46:14Z","receivedAt":"2007-05-01T21:46:14Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"OpenOffice.org is looking for a new SCM (Software Configuration \nManagement) tool, or at least was on Friday, 19 Jan 2007;\nsee: http://blogs.sun.com/GullFOSS/entry/openoffice_org_scm\n\nOne of the SCMs considered is Git. One of others is Subversion.\nThere is a functional git tree with the entire OOo history for testing \npurposes that can be found at: http://go-oo.org/git.\n\nWhat I am concerned about is some of git benchmark results at Git page \non OpenOffice.org wiki:\n  http://wiki.services.openoffice.org/wiki/Git#Comparison\nActually it is comparison with CVS and Subversion, although most \nbenchmarks are done only for git.\n\n\nIn 'Size of data on the server' git has CVS beat hands down: 1.3G vs \n8.5G for sources, 591M vs 1.1G for third party. I think it is similar\nfor Subversion. I hope that repository is fully packed: IIRC the Mozilla\nCVS repository import was about 0.6GB pack file, not 1.3GB.\n\n\nThe problem is with 'Size of checkout': to start working in repository\none needs 1.4G (sources) and 98M (third party) for CVS checkout (it is\n1.5G for sources for Subversion checkout). Ordinary for distributed SCM\nyou would need size of repository + size of sources (working area), \nwhich is 2.8G for sources and 688M for third party stuff files you can \nhack on + the history]. This makes some prefer to go centralized SCM \nroute, i.e. Subversion as replacement for CVS (+ CWS, ChildWorkSpace).\n\nWhat might help here is splitting repository into current (e.g. from\nOOo 2.0) and historical part, and / or using shallow clone. Implementing \npartial checkouts, i.e. checking out only part of working area (and \nusing 'theirs' strategy for merging not-checked-out part for merges) \nwould help. Splitting repository into submodules, and submodule\nsupport -- it depends on organization of OOo sources, would certainly \nhelp for third party stuff repository.\n\n'Checkout time' (which should be renamed to 'Initial checkout time'),\nin which git also loses with 130 minutes (Linux, 2MBit DSL) [from \ngo-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from \ngo-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux, \n2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows, \n34Mbit Line) for Subversion, would also be helped by the above.\n\n\nWhat I'm really concerned about is branch switch and merging branches,\nwhen one of the branches is an old one (e.g. unxsplash branch), which \ntakes 3min (!) according to the benchmark. 13-25sec for commit is also \nbit long, but BRANCH SWITCHING which takes 3 MINUTES!? There is no \ncomparison benchmark for CVS or Subversion, though...\n\nComparison / benchmark lacks some crucial info, like what computer was \nused (CPU, RAM, HDD), what filesystem was used, git version etc. It \ndoes have commands used for tests (benchmarks).\n\nCould you confirm (or deny) those results? go-oo.org uses git 1.4.3.4;\nwas there some improvement or bugfix related to the speed of checkout?\n\n-- \nJakub Narebski\nShadeHawk on #git\nPoland\n"},{"id":"40863","messageId":"7v4pmw5gdk.fsf@assigned-by-dhcp.cox.net","threadId":"7934","inReplyTo":"200705012346.14997.jnareb@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-05-01T22:27:35Z","receivedAt":"2007-05-01T22:27:35Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> What might help here is splitting repository into current (e.g. from\n> OOo 2.0) and historical part, and / or using shallow clone.\n\nYes, depending on where you cut off and how reasonable the\nproject history is.\n\n> Implementing \n> partial checkouts, i.e. checking out only part of working area (and \n> using 'theirs' strategy for merging not-checked-out part for merges) \n> would help.\n\nPartial checkouts, perhaps, \"theirs\", NO.\n\nConsider that you are working on the tip with partial checkout.\nSomebody has a bugfix that is applicable to all of ancient, old,\nmaintenance and current codebase.  Naturally you would want the\nbugfix to be applied to ancient, merge it to old, and then\nmaintenance and then current (the last one is what you are\nworking on).\n\nWhat happens if you actually pull ancient when you are partially\nchecked out and use \"theirs\"?\n\n> Splitting repository into submodules, and submodule\n> support -- it depends on organization of OOo sources, would certainly \n> help for third party stuff repository.\n\nThis is probably the most sane way.\n"},{"id":"40882","messageId":"200705020955.04582.andyparkins@gmail.com","threadId":"7934","inReplyTo":"200705012346.14997.jnareb@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-05-02T08:55:02Z","receivedAt":"2007-05-02T08:55:02Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Tuesday 2007 May 01, Jakub Narebski wrote:\n\n> In 'Size of data on the server' git has CVS beat hands down: 1.3G vs\n> 8.5G for sources, 591M vs 1.1G for third party. I think it is similar\n> for Subversion. I hope that repository is fully packed: IIRC the Mozilla\n> CVS repository import was about 0.6GB pack file, not 1.3GB.\n\nI'm fairly sure it's not.  If so that would also affect the speed of \noperations wouldn't it?\n\nI also doubt the subversion checkout size - subversion keeps a pristine copy \nof the HEAD file - so a subversion checkout is usually over twice the size of \nthe source tree.\n\n> takes 3min (!) according to the benchmark. 13-25sec for commit is also\n> bit long\n\nI wonder if they are measuring the time for the generation of the commit \nmessage or something?  Or perhaps by using \"git-commit -a\" is causing a check \nof the whole tree for changed files?\n\n> Comparison / benchmark lacks some crucial info, like what computer was\n> used (CPU, RAM, HDD), what filesystem was used, git version etc. It\n> does have commands used for tests (benchmarks).\n\nI'd also like to see some of the numbers for the other systems, I tried to use \nsubversion with the linux kernel once and got fed up waiting for it to do \nanything.  I suspect the reason numbers aren't shown for the others is that \nthey haven't finished yet :-)\n\n> Could you confirm (or deny) those results? go-oo.org uses git 1.4.3.4;\n> was there some improvement or bugfix related to the speed of checkout?\n\nWasn't there a recent change that made repacking after a clone unnecessary?  \nThat would certainly reduce the checkout size.\n\n\n\nAndy\n\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"40885","messageId":"Pine.LNX.4.64.0705021046230.2425@reaper.quantumfyre.co.uk","threadId":"7934","inReplyTo":"200705020955.04582.andyparkins@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-05-02T09:51:34Z","receivedAt":"2007-05-02T09:51:34Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Wed, 2 May 2007, Andy Parkins wrote:\n\n> On Tuesday 2007 May 01, Jakub Narebski wrote:\n>\n>> In 'Size of data on the server' git has CVS beat hands down: 1.3G vs\n>> 8.5G for sources, 591M vs 1.1G for third party. I think it is similar\n>> for Subversion. I hope that repository is fully packed: IIRC the Mozilla\n>> CVS repository import was about 0.6GB pack file, not 1.3GB.\n>\n> I'm fairly sure it's not.  If so that would also affect the speed of\n> operations wouldn't it?\n\nA fully packed clone of the OOo git repo was indeed 1.3G, and the entrire \ncheckout + repo was indeed 8.5G (using git 1.5.1.2).\n\nTook about 46m to clone on a server with decent bandwith, ~5.5m user time, \n~1.5m system.\n\n>\n> I also doubt the subversion checkout size - subversion keeps a pristine copy\n> of the HEAD file - so a subversion checkout is usually over twice the size of\n> the source tree.\n>\n>> takes 3min (!) according to the benchmark. 13-25sec for commit is also\n>> bit long\n>\n> I wonder if they are measuring the time for the generation of the commit\n> message or something?  Or perhaps by using \"git-commit -a\" is causing a check\n> of the whole tree for changed files?\n>\n>> Comparison / benchmark lacks some crucial info, like what computer was\n>> used (CPU, RAM, HDD), what filesystem was used, git version etc. It\n>> does have commands used for tests (benchmarks).\n>\n> I'd also like to see some of the numbers for the other systems, I tried to use\n> subversion with the linux kernel once and got fed up waiting for it to do\n> anything.  I suspect the reason numbers aren't shown for the others is that\n> they haven't finished yet :-)\n>\n>> Could you confirm (or deny) those results? go-oo.org uses git 1.4.3.4;\n>> was there some improvement or bugfix related to the speed of checkout?\n>\n> Wasn't there a recent change that made repacking after a clone unnecessary?\n> That would certainly reduce the checkout size.\n\nNot from the numbers that are quoted it won't, they are fully packed \nsizes.\n\n-- \nJulian\n\n  ---\n\"Do you have blacks, too?\"\n\nGeorge W. Bush\nTo Brazilian president Fernando Cardoso\nNovember 8, 2001\nWashington, D.C.\n"},{"id":"40887","messageId":"Pine.LNX.4.64.0705020143460.4010@racer.site","threadId":"7934","inReplyTo":"200705012346.14997.jnareb@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-02T10:24:32Z","receivedAt":"2007-05-02T10:24:32Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 1 May 2007, Jakub Narebski wrote:\n\n> 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> in which git also loses with 130 minutes (Linux, 2MBit DSL) [from \n> go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from \n> go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux, \n> 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows, \n> 34Mbit Line) for Subversion, would also be helped by the above.\n\nFWIW I can confirm the number \"100min\".\n\nSomething I realized with pain is that the refs/ directory is 24MB big. \nYep. Really. They have 3464 heads and 2639 tags. I suspect that this is \nthe reason why.\n\nWill play with it.\n\nCiao,\nDscho\n"},{"id":"40888","messageId":"200705021158.04481.andyparkins@gmail.com","threadId":"7934","inReplyTo":"Pine.LNX.4.64.0705021046230.2425@reaper.quantumfyre.co.uk","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-05-02T10:58:03Z","receivedAt":"2007-05-02T10:58:03Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Wednesday 2007 May 02, Julian Phillips wrote:\n\n> A fully packed clone of the OOo git repo was indeed 1.3G, and the entrire\n> checkout + repo was indeed 8.5G (using git 1.5.1.2).\n\nI'm more confused now then.  I assumed the figures were accurate, but they \ncannot be:\n\n                               CVS      git      SVN\nSize of data on the server     8.5G     1.3G     n/a\nSize of checkout               1.4G     2.8G     1.5G\n\nI don't doubt the 1.3G on the server - and assume that is fully packed.  The \ncheckout sizes are suspicious though.  Is that 2.8G packed?\n - If it is, then we can deduce that this is a repo+source size, since the\n   server is packed size+0 therefore the size of the source tree is\n    2.8G - 1.3G = 1.5G\n   In which case the other figures are wrong:\n    - CVS checkout is 1.4G - impossible, the source tree is 1.5G. And where is\n      the overhead of the CVS directories which would make it more than 1.5G?\n    - SVN checkout overhead is always _at least_ the size of the source tree \n      because it keeps a pristine copy of HEAD.  If the source tree is 1.5G,\n      then this figure should be at least 3G.\n - If it is not, then we're back to \"I don't believe that git was packed\"\n\nSomething smells fishy here - either the source tree size is included in some, \nbut not in others or the git repository wasn't packed.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"40890","messageId":"8fe92b430705020433v7ae5c117qdefccc791cd07fff@mail.gmail.com","threadId":"7934","inReplyTo":"Pine.LNX.4.64.0705020143460.4010@racer.site","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-02T11:33:55Z","receivedAt":"2007-05-02T11:33:55Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Hi!\n\nOn 5/2/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Tue, 1 May 2007, Jakub Narebski wrote:\n>\n> > 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> > in which git also loses with 130 minutes (Linux, 2MBit DSL) [from\n> > go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from\n> > go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux,\n> > 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows,\n> > 34Mbit Line) for Subversion, would also be helped by the above.\n>\n> FWIW I can confirm the number \"100min\".\n>\n> Something I realized with pain is that the refs/ directory is 24MB big.\n> Yep. Really. They have 3464 heads and 2639 tags. I suspect that this is\n> the reason why.\n\nThen packed refs would certainly help with speed and a bit with size.\n\n-- \nJakub Narebski\n"},{"id":"299255","messageId":"200705021624.25560.kendy@suse.cz","threadId":"7934","inReplyTo":"200705012346.14997.jnareb@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2007-05-02T14:24:24Z","receivedAt":"2007-05-02T14:24:24Z","isPatch":false,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Jakub,\n\nOn Tuesday 01 May 2007 23:46, Jakub Narebski wrote:\n\n> OpenOffice.org is looking for a new SCM (Software Configuration\n> Management) tool, or at least was on Friday, 19 Jan 2007;\n> see: http://blogs.sun.com/GullFOSS/entry/openoffice_org_scm\n>\n> One of the SCMs considered is Git. One of others is Subversion.\n> There is a functional git tree with the entire OOo history for testing\n> purposes that can be found at: http://go-oo.org/git.\n>\n> What I am concerned about is some of git benchmark results at Git page\n> on OpenOffice.org wiki:\n>   http://wiki.services.openoffice.org/wiki/Git#Comparison\n> Actually it is comparison with CVS and Subversion, although most\n> benchmarks are done only for git.\n\nI did the git numbers, so if they are wrong - blame me :-)  I am also curious\nabout the SVN numbers, because the SVN conversion [from my point of view]\ncheats a lot.  From what I know, it does not contain the historical branches\n(yes, the >3000 of them that are in the git tree), and if I understood that\ncorrectly, instead of history in the branches, they commit just\n'integration commits' [one commit for all the changes in the branch] which\nbreaks 'svn blame' completely.\n\nUnfortunately, I did not have a chance to try the SVN tree yet to see it\nmyself to prove this true or false :-(\n\n> In 'Size of data on the server' git has CVS beat hands down: 1.3G vs\n> 8.5G for sources, 591M vs 1.1G for third party. I think it is similar\n> for Subversion. I hope that repository is fully packed: IIRC the Mozilla\n> CVS repository import was about 0.6GB pack file, not 1.3GB.\n>\n> The problem is with 'Size of checkout': to start working in repository\n> one needs 1.4G (sources) and 98M (third party) for CVS checkout (it is\n> 1.5G for sources for Subversion checkout). Ordinary for distributed SCM\n> you would need size of repository + size of sources (working area),\n> which is 2.8G for sources and 688M for third party stuff files you can\n> hack on + the history]. This makes some prefer to go centralized SCM\n> route, i.e. Subversion as replacement for CVS (+ CWS, ChildWorkSpace).\n\nConsidering the size OOo needs for build (>8G without languages),\nthe ~1.4G overhead for history is very well bearable.  I am surprised about\nthe 100M overhead for SVN as well - from my experience it is usually about\nthe size of the project itself; but maybe they improved something in SVN\nin the meantime.\n\n> What might help here is splitting repository into current (e.g. from\n> OOo 2.0) and historical part,\n\nNo, I don't want this ;-)\n\n> and / or using shallow clone. Implementing \n> partial checkouts, i.e. checking out only part of working area (and\n> using 'theirs' strategy for merging not-checked-out part for merges)\n> would help. Splitting repository into submodules, and submodule\n> support -- it depends on organization of OOo sources, would certainly\n> help for third party stuff repository.\n\nWe should better split the OOo sources; it's a process that already started\n[UNO runtime environment vs. OOo without URE], and I proposed some more\nchanges already.\n\n> 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> in which git also loses with 130 minutes (Linux, 2MBit DSL) [from\n> go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from\n> go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux,\n> 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows,\n> 34Mbit Line) for Subversion, would also be helped by the above.\n\nGood point, and I already changed the page in the morning.  I also added the\ncheckout time that I got over a fast line [it was 44min].\n\n> What I'm really concerned about is branch switch and merging branches,\n> when one of the branches is an old one (e.g. unxsplash branch), which\n> takes 3min (!) according to the benchmark. 13-25sec for commit is also\n> bit long, but BRANCH SWITCHING which takes 3 MINUTES!? There is no\n> comparison benchmark for CVS or Subversion, though...\n\nI am really curious about the SVN tree.  As I said, I did not see it yet.\nThere is just some info about it here:\nhttp://wiki.services.openoffice.org/wiki/SVNMigration, but I cannot check it\nnow, the Wiki is down :-(\n\n> Comparison / benchmark lacks some crucial info, like what computer was\n> used (CPU, RAM, HDD), what filesystem was used, git version etc. It\n> does have commands used for tests (benchmarks).\n\nFor the git tests, it was:\n\nCPU: AMD Athlon(tm) 64 Processor 3200+\n\nRAM: 1G RAM\n\nDisk (info from bonnie):\n              ---Sequential Output (nosync)--- ---Sequential Input-- --Rnd Seek-\n              -Per Char- --Block--- -Rewrite-- -Per Char- --Block--- --04k (03)-\nMachine    MB K/sec %CPU K/sec %CPU K/sec %CPU K/sec %CPU K/sec %CPU   /sec %CPU\none    1*2000 37819 77.6 44296 16.8 16982  5.1 35203 63.9 45915  6.6  152.4  0.4\n\nFilesystem: ext3\n\n> Could you confirm (or deny) those results? go-oo.org uses git 1.4.3.4;\n> was there some improvement or bugfix related to the speed of checkout?\n\nRegards,\nJan\n\n"},{"id":"40897","messageId":"Pine.LNX.4.64.0705021523290.24218@reaper.quantumfyre.co.uk","threadId":"7934","inReplyTo":"200705021158.04481.andyparkins@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-05-02T14:28:55Z","receivedAt":"2007-05-02T14:28:55Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Wed, 2 May 2007, Andy Parkins wrote:\n\n> On Wednesday 2007 May 02, Julian Phillips wrote:\n>\n>> A fully packed clone of the OOo git repo was indeed 1.3G, and the entrire\n>> checkout + repo was indeed 8.5G (using git 1.5.1.2).\n\n\noops, meant 2.7G not 8.5G there ... sorry, was working from memory.\n\njp3@electron: ooo(unxsplash)>du -sh .git\n1.3G    .git\njp3@electron: ooo(unxsplash)>du -sh .\n2.7G    .\njp3@electron: ooo(unxsplash)>ls .git/objects/\ninfo  pack\n\n\n> I'm more confused now then.  I assumed the figures were accurate, but they\n> cannot be:\n>\n>                               CVS      git      SVN\n> Size of data on the server     8.5G     1.3G     n/a\n> Size of checkout               1.4G     2.8G     1.5G\n>\n> I don't doubt the 1.3G on the server - and assume that is fully packed.  The\n> checkout sizes are suspicious though.  Is that 2.8G packed?\n> - If it is, then we can deduce that this is a repo+source size, since the\n>   server is packed size+0 therefore the size of the source tree is\n>    2.8G - 1.3G = 1.5G\n\nthe difference between 2.7G and 2.8G may be due to filesystem difference?\n\n>   In which case the other figures are wrong:\n>    - CVS checkout is 1.4G - impossible, the source tree is 1.5G. And where is\n>      the overhead of the CVS directories which would make it more than 1.5G?\n>    - SVN checkout overhead is always _at least_ the size of the source tree\n>      because it keeps a pristine copy of HEAD.  If the source tree is 1.5G,\n>      then this figure should be at least 3G.\n\nI was wondering about the subversion figures too ...\n\n> - If it is not, then we're back to \"I don't believe that git was packed\"\n>\n> Something smells fishy here - either the source tree size is included in some,\n> but not in others or the git repository wasn't packed.\n\n1.3G is the packed size ...\n\njp3@electron: ooo(unxsplash)>ls -sh .git/objects/pack/\ntotal 1.3G\n  37M pack-87efcac9bcb117328e8a1b0c1b42c88c3603c5b7.idx\n1.2G pack-87efcac9bcb117328e8a1b0c1b42c88c3603c5b7.pack\n\n-- \nJulian\n\n  ---\nTo err is humor.\n"},{"id":"40898","messageId":"Pine.LNX.4.64.0705021630170.4015@racer.site","threadId":"7934","inReplyTo":"200705021624.25560.kendy@suse.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-02T14:35:17Z","receivedAt":"2007-05-02T14:35:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 2 May 2007, Jan Holesovsky wrote:\n\n> On Tuesday 01 May 2007 23:46, Jakub Narebski wrote:\n> \n> > What I am concerned about is some of git benchmark results at Git page\n> > on OpenOffice.org wiki:\n> >   http://wiki.services.openoffice.org/wiki/Git#Comparison\n> > Actually it is comparison with CVS and Subversion, although most\n> > benchmarks are done only for git.\n> \n> I did the git numbers, so if they are wrong - blame me :-)\n\nGood to have you here!\n\n> > 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> > in which git also loses with 130 minutes (Linux, 2MBit DSL) [from\n> > go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from\n> > go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux,\n> > 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows,\n> > 34Mbit Line) for Subversion, would also be helped by the above.\n> \n> Good point, and I already changed the page in the morning.  I also added the\n> checkout time that I got over a fast line [it was 44min].\n\nIt took me longer here, but the reason might be that my \"local\" repository \nis on NFS, due to quota on the machine.\n\n> > What I'm really concerned about is branch switch and merging branches,\n> > when one of the branches is an old one (e.g. unxsplash branch), which\n> > takes 3min (!) according to the benchmark. 13-25sec for commit is also\n> > bit long, but BRANCH SWITCHING which takes 3 MINUTES!? There is no\n> > comparison benchmark for CVS or Subversion, though...\n\nI imagine that might be related to the vast amount of remote branches. \nIIRC we do not pack them with git-gc, and ext3 is not that good with big \ndirectories (remember: 3464 branches!).\n\nMaybe oprofile knows a bit more where the hotspots are.\n\nCiao,\nDscho\n"},{"id":"40899","messageId":"200705021637.02057.kendy@suse.cz","threadId":"7934","inReplyTo":"200705021158.04481.andyparkins@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2007-05-02T14:37:01Z","receivedAt":"2007-05-02T14:37:01Z","isPatch":false,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Andy,\n\nOn Wednesday 02 May 2007 12:58, Andy Parkins wrote:\n\n> On Wednesday 2007 May 02, Julian Phillips wrote:\n> > A fully packed clone of the OOo git repo was indeed 1.3G, and the entrire\n> > checkout + repo was indeed 8.5G (using git 1.5.1.2).\n>\n> I'm more confused now then.  I assumed the figures were accurate, but they\n> cannot be:\n>\n>                                CVS      git      SVN\n> Size of data on the server     8.5G     1.3G     n/a\n> Size of checkout               1.4G     2.8G     1.5G\n>\n> I don't doubt the 1.3G on the server - and assume that is fully packed. \n> The checkout sizes are suspicious though.  Is that 2.8G packed?\n>  - If it is, then we can deduce that this is a repo+source size, since the\n>    server is packed size+0 therefore the size of the source tree is\n>     2.8G - 1.3G = 1.5G\n>    In which case the other figures are wrong:\n>     - CVS checkout is 1.4G - impossible, the source tree is 1.5G. And where\n> is the overhead of the CVS directories which would make it more than 1.5G?\n\nUnfortunately I don't have the _exact_ numbers here any more so I cannot prove \nit ;-) - but this is a rounding problem [CVS checkout is slightly more than \n1.4G].  Similarly, overhead of of CVS directories is 0 when we count in \ngigabytes.\n\n> - SVN checkout overhead is always _at least_ the size of the source tree\n> because it keeps a pristine copy of HEAD.  If the source tree is 1.5G, then\n> this figure should be at least 3G.\n\nYes, this surprises me as well.  I've heard about some improvements in the \nrecent SVN, but 0.1M sounds very small.\n\n>  - If it is not, then we're back to \"I don't believe that git was packed\"\n\nIt was, IIRC with 'git-repack -a -d -f'.\n\n> Something smells fishy here - either the source tree size is included in\n> some, but not in others or the git repository wasn't packed.\n\nAs I wrote, I am looking forward to seeing the SVN tree myself for further \ntesting.\n\nRegards,\nJan\n"},{"id":"40900","messageId":"200705021641.53199.kendy@suse.cz","threadId":"7934","inReplyTo":"Pine.LNX.4.64.0705020143460.4010@racer.site","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2007-05-02T14:41:52Z","receivedAt":"2007-05-02T14:41:52Z","isPatch":false,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Johannes,\n\nOn Wednesday 02 May 2007 12:24, Johannes Schindelin wrote:\n\n> On Tue, 1 May 2007, Jakub Narebski wrote:\n> > 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> > in which git also loses with 130 minutes (Linux, 2MBit DSL) [from\n> > go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from\n> > go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux,\n> > 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows,\n> > 34Mbit Line) for Subversion, would also be helped by the above.\n>\n> FWIW I can confirm the number \"100min\".\n>\n> Something I realized with pain is that the refs/ directory is 24MB big.\n> Yep. Really. They have 3464 heads and 2639 tags. I suspect that this is\n> the reason why.\n\nI should probably produce even a tree where would be the merged branches \ndeleted, right...\n\nRegards,\nJan\n"},{"id":"40901","messageId":"Pine.LNX.4.64.0705021653230.4015@racer.site","threadId":"7934","inReplyTo":"8fe92b430705020433v7ae5c117qdefccc791cd07fff@mail.gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-02T14:55:21Z","receivedAt":"2007-05-02T14:55:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 2 May 2007, Jakub Narebski wrote:\n\n> On 5/2/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > On Tue, 1 May 2007, Jakub Narebski wrote:\n> > \n> > > 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> > > in which git also loses with 130 minutes (Linux, 2MBit DSL) [from\n> > > go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from\n> > > go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux,\n> > > 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows,\n> > > 34Mbit Line) for Subversion, would also be helped by the above.\n> > \n> > FWIW I can confirm the number \"100min\".\n> > \n> > Something I realized with pain is that the refs/ directory is 24MB big.\n> > Yep. Really. They have 3464 heads and 2639 tags. I suspect that this is\n> > the reason why.\n> \n> Then packed refs would certainly help with speed and a bit with size.\n\nIndeed for size: du -h reported 11 megabyte for the tags directory. After \npacking them, a 265 kilobyte file is left. Of course, git-show-ref now \nbecomes a speed demon again.\n\nCiao,\nDscho\n"},{"id":"40902","messageId":"200705021630.16792.andyparkins@gmail.com","threadId":"7934","inReplyTo":"Pine.LNX.4.64.0705021523290.24218@reaper.quantumfyre.co.uk","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-05-02T15:30:14Z","receivedAt":"2007-05-02T15:30:14Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Wednesday 2007 May 02, Julian Phillips wrote:\n\n> oops, meant 2.7G not 8.5G there ... sorry, was working from memory.\n\nNot a problem.  That fixes one ambiguity:\n  2.7G - 1.3G = 1.4G \nWhich is the same as the CVS checkout size.  Both the CVS and git figures are \nnow consistent:\n                                CVS      git      SVN\nSize of data on the server     8.5G     1.3G      n/a\nSize of checkout               1.4G     2.7G     1.5G\nOverhead in checkout             0G     1.3G     0.1G\n\nSo that only leaves the subversion number as being suspicious.\n\n> the difference between 2.7G and 2.8G may be due to filesystem difference?\n\nCould be I suppose.  Although, in that case CVS should have suffered the same \nbecause the disparity was in the source tree size.  Packed git shouldn't \nsuffer any filesystem overhead (relatively) because the majority of it's \nspace is taken up by one large pack file (which of course only suffers file \nsystem overhead once).\n\n> I was wondering about the subversion figures too ...\n\nI've just checked using subversion 1.4.2 and the .svn/text-base/*.svn-base \nfiles are all uncompressed copies of the working tree files.  Doesn't look \nlike anythings changed in the pristine copy department.\n\n> jp3@electron: ooo(unxsplash)>ls -sh .git/objects/pack/\n> total 1.3G\n>   37M pack-87efcac9bcb117328e8a1b0c1b42c88c3603c5b7.idx\n> 1.2G pack-87efcac9bcb117328e8a1b0c1b42c88c3603c5b7.pack\n\nThanks for your help.  It's all looking more consistent to me now; only the \nsubversion figures seem wrong.\n\nI wonder when they're going to get timing numbers for the non-git systems.  \nThat must be a monster of a repository for them to deal with.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"40904","messageId":"200705021633.14428.andyparkins@gmail.com","threadId":"7934","inReplyTo":"200705021637.02057.kendy@suse.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Andy Parkins","fromEmail":"andyparkins@gmail.com","sentAt":"2007-05-02T15:33:11Z","receivedAt":"2007-05-02T15:33:11Z","isPatch":false,"sender":{"key":"andyparkins@gmail.com","avatar":null},"body":"On Wednesday 2007 May 02, Jan Holesovsky wrote:\n\n> Unfortunately I don't have the _exact_ numbers here any more so I cannot\n> prove it ;-) - but this is a rounding problem [CVS checkout is slightly\n> more than 1.4G].  Similarly, overhead of of CVS directories is 0 when we\n> count in gigabytes.\n\n0.1G would have been an awfully big rounding error.  Regardless, Julian has \nput me right on that - the git checked out size was actually 2.7GB - this \nthen lines up with the CVS figures.\n\n> > - SVN checkout overhead is always _at least_ the size of the source tree\n> > because it keeps a pristine copy of HEAD.  If the source tree is 1.5G,\n> > then this figure should be at least 3G.\n>\n> Yes, this surprises me as well.  I've heard about some improvements in the\n> recent SVN, but 0.1M sounds very small.\n\nVery much so - I've tried with a 1.4.2 and my own small repository and the \npristine copies are stored uncompressed as always.  0.1G now sounds plain \nwrong.  Maybe there are some switches I should be using to svn checkout.\n\n\nAndy\n-- \nDr Andy Parkins, M Eng (hons), MIET\nandyparkins@gmail.com\n"},{"id":"40905","messageId":"20070502161515.GC4489@pasky.or.cz","threadId":"7934","inReplyTo":"200705021624.25560.kendy@suse.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-05-02T16:15:15Z","receivedAt":"2007-05-02T16:15:15Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Wed, May 02, 2007 at 04:24:24PM CEST, Jan Holesovsky wrote:\n> > What might help here is splitting repository into current (e.g. from\n> > OOo 2.0) and historical part,\n> \n> No, I don't want this ;-)\n\nAre you sure? Using the graft mechanism, Git can make this very easy and\nalmost transparent for the user - when he clones he gets no history but\nhe can use say some simple vendor-provided script to download the\nhistorical packfile and graft it to the 'current' tree. After that, the\ngraft acts completely transparently and it 'seems' like the history\ngoes on continuously from OOo prehistory up to the latest commit.\n\nBesides, in case you discover a year later that the conversion was\nbroken in some places etc., you can just fix this, re-run the conversion\nand simply regraft your history to point at the 'new' historical commit,\nwithout affecting your current development and commit ids at all. For\nthis reason alone, I'd seriously consider grafting history separately\nwhen migrating any non-trivial project from other SCM to Git.\n\nThen again, due to the sheer tree sizes etc., I'm not sure how much\nwould throwing the history away actually reduce the packfile size.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"40907","messageId":"Pine.LNX.4.64.0705021822100.4015@racer.site","threadId":"7934","inReplyTo":"200705021641.53199.kendy@suse.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-02T16:24:12Z","receivedAt":"2007-05-02T16:24:12Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 2 May 2007, Jan Holesovsky wrote:\n\n> On Wednesday 02 May 2007 12:24, Johannes Schindelin wrote:\n> \n> > On Tue, 1 May 2007, Jakub Narebski wrote:\n> > > 'Checkout time' (which should be renamed to 'Initial checkout time'),\n> > > in which git also loses with 130 minutes (Linux, 2MBit DSL) [from\n> > > go-oo.org], 100min (Linux, 2MBit DSL, Wireless, no proxy) [from\n> > > go-oo.org] versus 117 minutes (Linux, 2MBit DSL), 26 minutes (Linux,\n> > > 2MBit DSL, with compression (-z 6)) for CVS, and  60 Minutes (Windows,\n> > > 34Mbit Line) for Subversion, would also be helped by the above.\n> >\n> > FWIW I can confirm the number \"100min\".\n> >\n> > Something I realized with pain is that the refs/ directory is 24MB big.\n> > Yep. Really. They have 3464 heads and 2639 tags. I suspect that this is\n> > the reason why.\n> \n> I should probably produce even a tree where would be the merged branches \n> deleted, right...\n\nFWIW, I just deleted all branches except for one, packed the tags, and did \na local clone (via NFS, urgh) _without_ checking the files out.\n\nNow it takes 25 minutes vs 50 minutes before (in an _extremely_ \nunscientific test, mind you).\n\nSo, this issue is worth looking at, probably.\n\nCiao,\nDscho\n"},{"id":"40908","messageId":"200705021827.51335.kendy@suse.cz","threadId":"7934","inReplyTo":"20070502161515.GC4489@pasky.or.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2007-05-02T16:27:51Z","receivedAt":"2007-05-02T16:27:51Z","isPatch":false,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Pasky,\n\nOn Wednesday 02 May 2007 18:15, Petr Baudis wrote:\n\n> On Wed, May 02, 2007 at 04:24:24PM CEST, Jan Holesovsky wrote:\n> > > What might help here is splitting repository into current (e.g. from\n> > > OOo 2.0) and historical part,\n> >\n> > No, I don't want this ;-)\n>\n> Are you sure? Using the graft mechanism, Git can make this very easy and\n> almost transparent for the user - when he clones he gets no history but\n> he can use say some simple vendor-provided script to download the\n> historical packfile and graft it to the 'current' tree. After that, the\n> graft acts completely transparently and it 'seems' like the history\n> goes on continuously from OOo prehistory up to the latest commit.\n\nInteresting, I did not know that it is possible to do it so that it appears \ntransparently; this would be indeed a tremendous win - we could start nearly \nfrom scratch ;-)\n\nPlease - where could I find more info?  Like what does the script have to do, \netc.\n\n> Besides, in case you discover a year later that the conversion was\n> broken in some places etc., you can just fix this, re-run the conversion\n> and simply regraft your history to point at the 'new' historical commit,\n> without affecting your current development and commit ids at all. For\n> this reason alone, I'd seriously consider grafting history separately\n> when migrating any non-trivial project from other SCM to Git.\n>\n> Then again, due to the sheer tree sizes etc., I'm not sure how much\n> would throwing the history away actually reduce the packfile size.\n\nThanks a lot,\nJan\n"},{"id":"40911","messageId":"20070502163715.GD4489@pasky.or.cz","threadId":"7934","inReplyTo":"200705021827.51335.kendy@suse.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-05-02T16:37:15Z","receivedAt":"2007-05-02T16:37:15Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hi,\n\nOn Wed, May 02, 2007 at 06:27:51PM CEST, Jan Holesovsky wrote:\n> On Wednesday 02 May 2007 18:15, Petr Baudis wrote:\n> \n> > On Wed, May 02, 2007 at 04:24:24PM CEST, Jan Holesovsky wrote:\n> > > > What might help here is splitting repository into current (e.g. from\n> > > > OOo 2.0) and historical part,\n> > >\n> > > No, I don't want this ;-)\n> >\n> > Are you sure? Using the graft mechanism, Git can make this very easy and\n> > almost transparent for the user - when he clones he gets no history but\n> > he can use say some simple vendor-provided script to download the\n> > historical packfile and graft it to the 'current' tree. After that, the\n> > graft acts completely transparently and it 'seems' like the history\n> > goes on continuously from OOo prehistory up to the latest commit.\n> \n> Interesting, I did not know that it is possible to do it so that it appears \n> transparently; this would be indeed a tremendous win - we could start nearly \n> from scratch ;-)\n> \n> Please - where could I find more info?  Like what does the script have to do, \n> etc.\n\n  you can see an example script at\n\n\thttp://repo.or.cz/w/elinks.git?a=blob;f=contrib/grafthistory.sh\n\nand I have tried vainly few times to get a similar script to the kernel\ntoo\n\n\thttp://lists.zerezo.com/linux-kernel/msg6599002.html\n\nthat can use both wget and curl and will also download tag refs for the\nhistory.\n\n  The format of the grafts file itself (.git/info/grafts) is pretty\nsimple (just one-graft-per-line where you first say the commit id and\nthen the parent commit(s) to be drafted onto it), please see\nDocumentation/repository-layout.txt for details.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"40912","messageId":"20070502164821.GE4489@pasky.or.cz","threadId":"7934","inReplyTo":"20070502163715.GD4489@pasky.or.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-05-02T16:48:21Z","receivedAt":"2007-05-02T16:48:21Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hi,\n\nOn Wed, May 02, 2007 at 06:37:15PM CEST, Petr Baudis wrote:\n>   you can see an example script at\n> \n> \thttp://repo.or.cz/w/elinks.git?a=blob;f=contrib/grafthistory.sh\n\n  by the way, this script goes back to very ancient Git times, maybe by\nnow git-fetch could be convinced to do all the hard work for you.\nActually, maybe just something (totally untested) like\n\n\tgit remote add -f historical {http,git}://historical_repository_url\n\tcat <<EOF >>.git/info/grafts\n\t... the graft specs go here ...\n\tEOF\n\nmight work prefectly fine nowadays that git keeps the remote refs in a\nseparate namespace tidily. This way you don't have to care about all the\nmanual wgetting, ls-remote magic etc. The downside is that this is\navailable only since git-1.5.0 (Debian stable has older version; maybe\neven newer git version is required, I'm not sure).\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"40914","messageId":"Pine.LNX.4.64.0705021801120.20300@reaper.quantumfyre.co.uk","threadId":"7934","inReplyTo":"200705021630.16792.andyparkins@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-05-02T17:11:56Z","receivedAt":"2007-05-02T17:11:56Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Wed, 2 May 2007, Andy Parkins wrote:\n\n> On Wednesday 2007 May 02, Julian Phillips wrote:\n>\n>> oops, meant 2.7G not 8.5G there ... sorry, was working from memory.\n>\n> Not a problem.  That fixes one ambiguity:\n>  2.7G - 1.3G = 1.4G\n> Which is the same as the CVS checkout size.  Both the CVS and git figures are\n> now consistent:\n>                                CVS      git      SVN\n> Size of data on the server     8.5G     1.3G      n/a\n> Size of checkout               1.4G     2.7G     1.5G\n> Overhead in checkout             0G     1.3G     0.1G\n\nExcept that it's 2.8G, I forgot I had switched branch.  I switched to the \nunxsplash branch, and _that_ is 2.7G checked out.\n\n(du -s .) - (du -s .git) = 1.49G\n\n-- \nJulian\n\n  ---\n\"Consider a spherical bear, in simple harmonic motion...\"\n \t\t-- Professor in the UCB physics department\n"},{"id":"40916","messageId":"7vodl33znz.fsf@assigned-by-dhcp.cox.net","threadId":"7934","inReplyTo":"200705021158.04481.andyparkins@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-05-02T17:26:08Z","receivedAt":"2007-05-02T17:26:08Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andy Parkins <andyparkins@gmail.com> writes:\n\n>     - SVN checkout overhead is always _at least_ the size of the source tree \n>       because it keeps a pristine copy of HEAD.  If the source tree is 1.5G,\n>       then this figure should be at least 3G.\n\nCould it be that there is a mode in svn checkout that allows\npristine to be hardlinked to the working tree copies?  It\nrequires an editor that can be told to break hardlinks when\nmaking modifications (and the user obviously needs to know about\nit), but to save 1.5G it is worth it and if _I_ were hacking on\nSVN that would be an obvious optimization to add.\n"},{"id":"40927","messageId":"200705030130.44018.jnareb@gmail.com","threadId":"7934","inReplyTo":"200705021624.25560.kendy@suse.cz","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-02T23:30:43Z","receivedAt":"2007-05-02T23:30:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jan Holesovsky wrote:\n> On Tuesday 01 May 2007 23:46, Jakub Narebski wrote:\n> \n>> What I am concerned about is some of git benchmark results at Git page\n>> on OpenOffice.org wiki:\n>>   http://wiki.services.openoffice.org/wiki/Git#Comparison\n\n>> The problem is with 'Size of checkout': to start working in repository\n>> one needs 1.4G (sources) and 98M (third party) for CVS checkout (it is\n>> 1.5G for sources for Subversion checkout). Ordinary for distributed SCM\n>> you would need size of repository + size of sources (working area),\n>> which is 2.8G for sources and 688M for third party stuff files you can\n>> hack on + the history]. This makes some prefer to go centralized SCM\n>> route, i.e. Subversion as replacement for CVS (+ CWS, ChildWorkSpace).\n> \n> Considering the size OOo needs for build (>8G without languages),\n> the ~1.4G overhead for history is very well bearable.  I am surprised about\n> the 100M overhead for SVN as well - from my experience it is usually about\n> the size of the project itself; but maybe they improved something in SVN\n> in the meantime.\n\nI think the supposition that SVN uses hardlinks for pristine copy\nof sources (HEAD version) seems probable; then there it is 100M overhead\nplus size of changed files, and of course this tricks works only on\nfilesystems which support hardlinks, and assumes either hardlinks being\nCOW-links (copy-on-write) or editor behaving.\n \n>> What might help here is splitting repository into current (e.g. from\n>> OOo 2.0) and historical part,\n> \n> No, I don't want this ;-)\n\nI forgot to add there is possible to graft historical repository to the\ncurrent work repository, resulting in full history available. For example\nLinux kernel repository has backported from BK historical repository, and\nthere is grafts file which connect those two repositories.\n\n>> and / or using shallow clone. \n\ngit-clone(1):\n\n--depth <depth>::\n        Create a 'shallow' clone with a history truncated to the\n        specified number of revs.  A shallow repository has\n        number of limitations (you cannot clone or fetch from\n        it, nor push from nor into it), but is adequate if you\n        want to only look at near the tip of a large project\n        with a long history, and would want to send in a fixes\n        as patches.\n\nIt is possible that those limitations will be lifted in the future\n(if possible), so there is alternate possibility to reduce needed\ndisk space for git checkout. But certainly this is not for everybody.\n\n>> Implementing  \n>> partial checkouts, i.e. checking out only part of working area (...)\n\nThe problem with implementing this feature (you can do partial checkout\nusing low level commands, but this feature is not implemented [yet?]\nper se) is with doing merge on part which is not checked out. Might\nnot be a problem for OOo; but this might be also not needed for OOo.\nSometimes submodules are better, sometimes partial checkout is the\nonly way: see below.\n\n>> Splitting repository into submodules, and submodule \n>> support -- it depends on organization of OOo sources, would certainly\n>> help for third party stuff repository.\n> \n> We should better split the OOo sources; it's a process that already started\n> [UNO runtime environment vs. OOo without URE], and I proposed some more\n> changes already.\n\nIn my opinion each submodule should be able to compile and test by\nitself. You can go X.Org route with splitting sources into modules...\nor you can make use of the new submodules support (currently plumbing\nlevel, i.e. low level commands), aka. gitlinks.\n\nThe submodules support makes it possible to split sources into\nindependent modules (parts), which can be developed independently,\nand which you can download (clone, fetch) or not, while making it\npossible to bind it all together into one superproject.\n\nSee (somewhat not up to date) http://git.or.cz/gitwiki/SubprojectSupport\npage on git wiki.\n\n>> What I'm really concerned about is branch switch and merging branches,\n>> when one of the branches is an old one (e.g. unxsplash branch), which\n>> takes 3min (!) according to the benchmark. 13-25sec for commit is also\n>> bit long, but BRANCH SWITCHING which takes 3 MINUTES!? There is no\n>> comparison benchmark for CVS or Subversion, though...\n\nBy the way, the time to switch branch should be proportional to number\nof changed files, which you can get with \"git diff --summary unxsplash\nHEAD\". Or to be more realistic to checkout some old version\n(some old tag), as usually branches which got merged in are deleted\n(or even never got published). For example when bisecting some bug:\nSubversion doesn't have bisect, does it?\n\nI wonder if running \"git pack-refs\" would help this benchmark...\n\n-- \nJakub Narebski\nPoland\n"},{"id":"40948","messageId":"87lkg61j99.fsf@mid.deneb.enyo.de","threadId":"7934","inReplyTo":"200705012346.14997.jnareb@gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2007-05-03T07:03:30Z","receivedAt":"2007-05-03T07:03:30Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* Jakub Narebski:\n\n> The problem is with 'Size of checkout': to start working in repository\n> one needs 1.4G (sources) and 98M (third party) for CVS checkout (it is\n> 1.5G for sources for Subversion checkout).\n\nThe text bases for Subversion really should take another 1.4 GiB.\nAs a result, Subversion should be closer to 3 GiB.\n\n> What might help here is splitting repository into current (e.g. from\n> OOo 2.0) and historical part, and / or using shallow clone.\n\nYou could also split along project boundaries, but this is probably\ntoo political.\n\n> What I'm really concerned about is branch switch and merging branches,\n> when one of the branches is an old one (e.g. unxsplash branch), which \n> takes 3min (!) according to the benchmark. 13-25sec for commit is also \n> bit long, but BRANCH SWITCHING which takes 3 MINUTES!?\n\nIIRC, GIT accesses every file in the tree, not just the ones that need\nupdating.  How many files were actually updated when you changed\nbranches in your experiment?\n"},{"id":"40952","messageId":"Pine.LNX.4.64.0705031131410.4015@racer.site","threadId":"7934","inReplyTo":"87lkg61j99.fsf@mid.deneb.enyo.de","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-03T09:33:21Z","receivedAt":"2007-05-03T09:33:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 3 May 2007, Florian Weimer wrote:\n\n> * Jakub Narebski:\n> \n> > What I'm really concerned about is branch switch and merging branches, \n> > when one of the branches is an old one (e.g. unxsplash branch), which \n> > takes 3min (!) according to the benchmark. 13-25sec for commit is also \n> > bit long, but BRANCH SWITCHING which takes 3 MINUTES!?\n> \n> IIRC, GIT accesses every file in the tree, not just the ones that need\n> updating.  How many files were actually updated when you changed\n> branches in your experiment?\n\nNo. Git does not access every file, but rather all stats. That is a huge \ndifference. And it should not take _that_ long for ~64000 files. Granted, \nit will cause a substantial delay, but not in the range of minutes.\n\nCiao,\nDscho\n"},{"id":"40953","messageId":"200705031216.19817.robin.rosenberg.lists@dewire.com","threadId":"7934","inReplyTo":"Pine.LNX.4.64.0705031131410.4015@racer.site","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-05-03T10:16:05Z","receivedAt":"2007-05-03T10:16:05Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdag 03 maj 2007 skrev Johannes Schindelin:\n> Hi,\n> \n> On Thu, 3 May 2007, Florian Weimer wrote:\n> \n> > * Jakub Narebski:\n> > \n> > > What I'm really concerned about is branch switch and merging branches, \n> > > when one of the branches is an old one (e.g. unxsplash branch), which \n> > > takes 3min (!) according to the benchmark. 13-25sec for commit is also \n> > > bit long, but BRANCH SWITCHING which takes 3 MINUTES!?\n> > \n> > IIRC, GIT accesses every file in the tree, not just the ones that need\n> > updating.  How many files were actually updated when you changed\n> > branches in your experiment?\n> \n> No. Git does not access every file, but rather all stats. That is a huge \n> difference. And it should not take _that_ long for ~64000 files. Granted, \n> it will cause a substantial delay, but not in the range of minutes.\n\nIt's worse... On my laptop the switch took ~ten minutes, not three. \nA diff --stat takes over six minutes!! For reference, dd:in the pack \nfile with my disk takes ~50 seconds.\n\nThe reason is simple. I have a lousy one gigabyte RAM only, while \ngit wants 1.7GB virtual to do the diff-stat.  and 800 MB resident. The swap is having a party,\n\n$ vmstat 1\nprocs -----------memory---------- ---swap-- -----io---- -system-- ----cpu----\n r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa\n 1  2 1861632  14108    428 126816   70  347   605   647  594 1041 11  2 74 13\n 0  2 1861204  12096    420 125724 3096    8  3096    24  625 1171  5  1  0 94\n 0  2 1860896  18972    404 115836 3524  292  3524   292  671 1474  7  4  0 89\n 0  2 1860820  18668    364 113736 3556  784  3556   784  669 1384  7  5  0 88\n 0  3 1860420  19692    300 109904 3008  180  3156   180  684 1325  8  5  0 87\n 0  3 1860184  18560    300 108596 3316  232  3396   232  643 1246  8  4  0 88\n 0  2 1859856  21808    292 103744 2108   32  2356    32  637 1319  9  1  0 90\n\n-- robin\n"},{"id":"40954","messageId":"46a038f90705030348o260fbe6cwc92d07778269c937@mail.gmail.com","threadId":"7934","inReplyTo":"200705031216.19817.robin.rosenberg.lists@dewire.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-05-03T10:48:40Z","receivedAt":"2007-05-03T10:48:40Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"[resend - correcting a couple of typos and addressing git@vger\ncorrectly - apologies]\n\nOn 5/3/07, Robin Rosenberg <robin.rosenberg.lists@dewire.com> wrote:\n> The reason is simple. I have a lousy one gigabyte RAM only, while\n> git wants 1.7GB virtual to do the diff-stat.  and 800 MB resident. The swap is having a party,\n\nThat is true, unfortunately. git will fly if it can fit its working\nset plus the kernel stat cache for your working tree in memory. And\nthe underlying assumption is that for large trees you'll have gobs of\nRAM. If things don't fit, it does get rather slow...\n\nBut... just to put things in perspective, how long does it take to\n*compile* that checkout on that same laptop. I remember reading\ninstructions to the tune of \"don't even try to compile this with less\nthan 4GB RAM, a couple of CPUs and 12hs\". Those were for the OSX build\nIIRC.\n\nAh - it's moved to the general instructions: \"Building OOo takes some\ntime (approx 10-12 hours on standard desktop PC) \":\nhttp://wiki.services.openoffice.org/wiki/Building_OpenOffice.org#Starting_the_real_build\n\nSo I don't think anyone working on projects the size of the kernel or\nOO.org is going to be happy with 1GB RAM.\n\ncheers\n\n\nm\n"},{"id":"40956","messageId":"200705031351.40548.kendy@suse.cz","threadId":"7934","inReplyTo":"200705030130.44018.jnareb@gmail.com","subject":"Re: [tools-dev] Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jan Holesovsky","fromEmail":"kendy@suse.cz","sentAt":"2007-05-03T11:51:40Z","receivedAt":"2007-05-03T11:51:40Z","isPatch":false,"sender":{"key":"kendy@suse.cz","avatar":null},"body":"Hi Jakub,\n\nOn Thursday 03 May 2007 01:30, Jakub Narebski wrote:\n\n> >> What might help here is splitting repository into current (e.g. from\n> >> OOo 2.0) and historical part,\n> >\n> > No, I don't want this ;-)\n>\n> I forgot to add there is possible to graft historical repository to the\n> current work repository, resulting in full history available. For example\n> Linux kernel repository has backported from BK historical repository, and\n> there is grafts file which connect those two repositories.\n\nYes, grafting sounds really very promising! - I did not know about it.\n\n> >> and / or using shallow clone.\n>\n> git-clone(1):\n>\n> --depth <depth>::\n>         Create a 'shallow' clone with a history truncated to the\n>         specified number of revs.  A shallow repository has\n>         number of limitations (you cannot clone or fetch from\n>         it, nor push from nor into it), but is adequate if you\n>         want to only look at near the tip of a large project\n>         with a long history, and would want to send in a fixes\n>         as patches.\n>\n> It is possible that those limitations will be lifted in the future\n> (if possible), so there is alternate possibility to reduce needed\n> disk space for git checkout. But certainly this is not for everybody.\n\nIt's probably too tight limitation for regular developers; for random hackers \ncontributing a patch or two it could be a choice, right.\n\n> > We should better split the OOo sources; it's a process that already\n> > started [UNO runtime environment vs. OOo without URE], and I proposed\n> > some more changes already.\n>\n> In my opinion each submodule should be able to compile and test by\n> itself. You can go X.Org route with splitting sources into modules...\n\nIndeed, this is the case of URE - it is supposed to run by separately & be \nused even by other projects than OOo.\n\n> or you can make use of the new submodules support (currently plumbing\n> level, i.e. low level commands), aka. gitlinks.\n\nAnd this would be interesting for the translations, I guess...\n\n> The submodules support makes it possible to split sources into\n> independent modules (parts), which can be developed independently,\n> and which you can download (clone, fetch) or not, while making it\n> possible to bind it all together into one superproject.\n>\n> See (somewhat not up to date) http://git.or.cz/gitwiki/SubprojectSupport\n> page on git wiki.\n\n... but will have a better look; thanks for the pointer!\n\n> Subversion doesn't have bisect, does it?\n\n>From what I know, it does not.\n\nThank you and others for all the input!\n\nLast question: what is the status of the Win32 support?  I got a full clone \nusing the Cygwin git 1.5.0 [it took 6hrs 20min on a Xen virtual machine; I \nhave to try it with real hardware], MinGW version did not work for me too \nwell :-(  Are there any other options?  Is \nhttp://git.or.cz/gitwiki/WindowsInstall up-to-date?\n\nRegards,\nJan\n"},{"id":"40959","messageId":"81b0412b0705030554jd8628d1tde8f5c1135900c95@mail.gmail.com","threadId":"7934","inReplyTo":"200705031351.40548.kendy@suse.cz","subject":"Re: [tools-dev] Re: Git benchmarks at OpenOffice.org wiki","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2007-05-03T12:54:25Z","receivedAt":"2007-05-03T12:54:25Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 5/3/07, Jan Holesovsky <kendy@suse.cz> wrote:\n> Last question: what is the status of the Win32 support?\n\nIt kind of works. Performance is horrible, but still better\nthan almost everything comparable (and there isn't anything\ncomparable). You have to be very careful not to push it\n(them, actually: cygwin and windows) too hard: it is quick\nto fall over taking down the whole machine with it (yes,\navoid Ctrl-C at all costs).\nThe repos have always been recoverable for me, though.\n\n> I got a full clone using the Cygwin git 1.5.0 [it took 6hrs 20min\n> on a Xen virtual machine; I have to try it with real hardware],\n> MinGW version did not work for me too well :-(\n>  Are there any other options?\n\nAvoid Win32 if possible, work somewhere in a sane environment,\nusing windows for testing, if you have to.\n\n> Is http://git.or.cz/gitwiki/WindowsInstall up-to-date?\n\nYes.\n"},{"id":"40964","messageId":"4639FC5D.8F24E45E@eudaptics.com","threadId":"7934","inReplyTo":"200705031351.40548.kendy@suse.cz","subject":"Re: [tools-dev] Re: Git benchmarks at OpenOffice.org wiki","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2007-05-03T15:14:37Z","receivedAt":"2007-05-03T15:14:37Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Jan Holesovsky wrote:\n> Last question: what is the status of the Win32 support?  I got a full clone\n> using the Cygwin git 1.5.0 [it took 6hrs 20min on a Xen virtual machine; I\n> have to try it with real hardware], MinGW version did not work for me too\n> well :-(\n\nI'd really like to hear what you exactly mean by \"MinGW version did not\nwork too well\". I use it in production. (But then, I'm the one who\npublishes the MinGW port ;)\n\n-- Hannes\n"},{"id":"40995","messageId":"f1drdk$7uj$1@sea.gmane.org","threadId":"7934","inReplyTo":"200705031216.19817.robin.rosenberg.lists@dewire.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-03T23:36:27Z","receivedAt":"2007-05-03T23:36:27Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Robin Rosenberg wrote:\n> torsdag 03 maj 2007 skrev Johannes Schindelin:\n>> On Thu, 3 May 2007, Florian Weimer wrote:\n>>> * Jakub Narebski:\n>>> \n>>>> What I'm really concerned about is branch switch and merging branches, \n>>>> when one of the branches is an old one (e.g. unxsplash branch), which \n>>>> takes 3min (!) according to the benchmark. 13-25sec for commit is also \n>>>> bit long, but BRANCH SWITCHING which takes 3 MINUTES!?\n>>> \n>>> IIRC, GIT accesses every file in the tree, not just the ones that need\n>>> updating.  How many files were actually updated when you changed\n>>> branches in your experiment?\n>> \n>> No. Git does not access every file, but rather all stats. That is a huge \n>> difference. And it should not take _that_ long for ~64000 files. Granted, \n>> it will cause a substantial delay, but not in the range of minutes.\n> \n> It's worse... On my laptop the switch took ~ten minutes, not three. \n> A diff --stat takes over six minutes!! For reference, dd:in the pack \n> file with my disk takes ~50 seconds.\n> \n> The reason is simple. I have a lousy one gigabyte RAM only, while \n> git wants 1.7GB virtual to do the diff-stat, and 800 MB resident.\n> The swap is having a party, \n\nThat is nice to know where the culprit is: 197000 files and 24000\ndirectories (132 projects), i.e. huge tree and not enogh memory.\nThis is yet another reason for splitting OOo repository into subprojects.\nI do wonder if git can be more conservative in memory usage...\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"41001","messageId":"200705040248.22443.jnareb@gmail.com","threadId":"7934","inReplyTo":"200705031351.40548.kendy@suse.cz","subject":"Re: [tools-dev] Re: Git benchmarks at OpenOffice.org wiki","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-04T00:48:22Z","receivedAt":"2007-05-04T00:48:22Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Thu, May 02, 2007, Jan Holesovsky wrote:\n> On Thursday 03 May 2007 01:30, Jakub Narebski wrote:\n\n>>> We should better split the OOo sources; it's a process that already\n>>> started [UNO runtime environment vs. OOo without URE], and I proposed\n>>> some more changes already.\n>>\n>> In my opinion each submodule should be able to compile and test by\n>> itself. You can go X.Org route with splitting sources into modules...\n> \n> Indeed, this is the case of URE - it is supposed to run by separately & be \n> used even by other projects than OOo.\n> \n>> or you can make use of the new submodules support (currently plumbing\n>> level, i.e. low level commands), aka. gitlinks.\n> \n> And this would be interesting for the translations, I guess...\n> \n>> The submodules support makes it possible to split sources into\n>> independent modules (parts), which can be developed independently,\n>> and which you can download (clone, fetch) or not, while making it\n>> possible to bind it all together into one superproject.\n\nBy the way, even without submodule support, which for now is plumbing\nlevel only, it would be possible to pull separate subprojects into main\nproject, like git repository does now with gitk repository, and with\ngit-gui repository. The latter is merged putting git-gui files in separate\ndirectory in git.git repository, via using 'subtree' merge strategy.\n\nSubmodules / subprojects are something similar to Subversion svn:externals\ndone right.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"41116","messageId":"alpine.LFD.0.98.0705042049330.3819@woody.linux-foundation.org","threadId":"7934","inReplyTo":"8fe92b430705020433v7ae5c117qdefccc791cd07fff@mail.gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-05-05T03:56:08Z","receivedAt":"2007-05-05T03:56:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 2 May 2007, Jakub Narebski wrote:\n> On 5/2/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > \n> > Something I realized with pain is that the refs/ directory is 24MB big.\n> > Yep. Really. They have 3464 heads and 2639 tags. I suspect that this is\n> > the reason why.\n> \n> Then packed refs would certainly help with speed and a bit with size.\n\nBtw, this reminds me: we really should start out clones with a fully \npacked set of refs. It seems stupid to get the refs in one go, and then \nexplode them into thousands of files.\n\nA trivial patch is to just do\n\n\tgit pack-refs --all --prune\n\nin the \"git-clone.sh\" script rather than force people to do it themselves, \nbut we really probably shouldn't have ever even unpacked them in the first \nplace. That is kind of stupid, but especially since that thing is written \nin shell, it's hard to do anything smarter.\n\nOf course, I don't know what the hell openoffice is doing with that many \nbranches and tags, but I guess it's a normal result of having used CVS/SVN \n- you want to tag every single merge you do, and all branches stay around \nforever, because you can never merge them back and get rid of them.\n\nIt's always sad to see the crap that is CVS, and how bad decisions in CVS \nend up resulting in pain downstream.\n\n\t\tLinus\n"},{"id":"41246","messageId":"200705062205.34814.robin.rosenberg.lists@dewire.com","threadId":"7934","inReplyTo":"46a038f90705030348o260fbe6cwc92d07778269c937@mail.gmail.com","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-05-06T20:05:30Z","receivedAt":"2007-05-06T20:05:30Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdag 03 maj 2007 skrev Martin Langhoff:\n> [resend - correcting a couple of typos and addressing git@vger\n> correctly - apologies]\n> \n> On 5/3/07, Robin Rosenberg <robin.rosenberg.lists@dewire.com> wrote:\n> > The reason is simple. I have a lousy one gigabyte RAM only, while\n> > git wants 1.7GB virtual to do the diff-stat.  and 800 MB resident. The swap is having a party,\n> \n> That is true, unfortunately. git will fly if it can fit its working\n> set plus the kernel stat cache for your working tree in memory. And\n> the underlying assumption is that for large trees you'll have gobs of\n> RAM. If things don't fit, it does get rather slow...\n> \n> But... just to put things in perspective, how long does it take to\n> *compile* that checkout on that same laptop. I remember reading\n> instructions to the tune of \"don't even try to compile this with less\n> than 4GB RAM, a couple of CPUs and 12hs\". Those were for the OSX build\n> IIRC.\nNo idea. I wouldn't try it without distcc and ccache anyway which makes the\ncapabilities of this particular machine less relevant. \n\n> Ah - it's moved to the general instructions: \"Building OOo takes some\n> time (approx 10-12 hours on standard desktop PC) \":\n> http://wiki.services.openoffice.org/wiki/Building_OpenOffice.org#Starting_the_real_build\n> \n> So I don't think anyone working on projects the size of the kernel or\n> OO.org is going to be happy with 1GB RAM.\n\nThe kernel 2.6 repo isn't in the same ball park wrt to size.  Hacking the kernel is quite fine\non this machine and even smaller though the first compile takes some time. Having more \nis always fun though. \n\nConsider another huge project like Eclipse. Similar operations take a loong time (not anywhere\nnear the eons that CVS need, but...)  and building Eclipse with 1GB i very reasonable so sheer\nproject size does not per se demand powerful computers. KDE is another huge project that\nis reasonable to build with 1GB. The first time is somewhat painful, but rebuilding is not.\n\n-- robin\n"},{"id":"41304","messageId":"7vk5vl6oum.fsf@assigned-by-dhcp.cox.net","threadId":"7934","inReplyTo":"alpine.LFD.0.98.0705042049330.3819@woody.linux-foundation.org","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-05-07T08:05:05Z","receivedAt":"2007-05-07T08:05:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> A trivial patch is to just do\n>\n> \tgit pack-refs --all --prune\n>\n> in the \"git-clone.sh\" script rather than force people to do it themselves, \n> but we really probably shouldn't have ever even unpacked them in the first \n> place. That is kind of stupid, but especially since that thing is written \n> in shell, it's hard to do anything smarter.\n\nIt is not just being in shell.\n\nAlthough I do agree that the initial clone is special, I would\nrather make clone just a thin wrapper to fetch that also happens\nto perform necessary initial setup.\n\nKeeping fetched and updated refs in core and write a packed refs\nout in one go in git-fetch--tool (and later, git-fetch all in C)\nwould be much simpler if we do not have to worry about existing\nrefs (aka \"git clone\" special case); I am not sure if packing\nrefs is desirable in general for incremental \"git-fetch\".\n"},{"id":"41342","messageId":"alpine.LFD.0.98.0705070821260.3802@woody.linux-foundation.org","threadId":"7934","inReplyTo":"7vk5vl6oum.fsf@assigned-by-dhcp.cox.net","subject":"Re: Git benchmarks at OpenOffice.org wiki","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-05-07T15:22:52Z","receivedAt":"2007-05-07T15:22:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 7 May 2007, Junio C Hamano wrote:\n> \n> Keeping fetched and updated refs in core and write a packed refs\n> out in one go in git-fetch--tool (and later, git-fetch all in C)\n> would be much simpler if we do not have to worry about existing\n> refs (aka \"git clone\" special case); I am not sure if packing\n> refs is desirable in general for incremental \"git-fetch\".\n\nFair enough. It's true that for the general case of \"git fetch\", it's much \nless obvious how to keep things packed.\n\nSo maybe the right thing really *is* to just add the\n\n\tgit pack-refs --all --prune\n\nto the git-clone wrapper.\n\n\t\tLinus\n"}]}