{"thread":{"id":"36670","subject":"Re: worlds slowest git repo- what to do?","startedAt":"2014-05-15T19:06:40Z","lastAt":"2014-05-17T01:49:32Z","messageCount":5,"participants":["Philip Oakley","Sam Vilain","Duy Nguyen","John Fisher"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"241766","messageId":"06A2490FC9BC4461A39B982D3C7C85F7@PhilipOakley","threadId":"36670","inReplyTo":"5374F7C6.5030205@gmail.com","subject":"Re: worlds slowest git repo- what to do?","fromName":"Philip Oakley","fromEmail":"philipoakley-7kbabnvhqfm@public.gmane.org","sentAt":"2014-05-15T19:06:40Z","receivedAt":"2014-05-15T19:06:40Z","isPatch":false,"sender":{"key":"philipoakley-7kbabnvhqfm@public.gmane.org","avatar":null},"body":"From: \"John Fisher\" <fishook2033-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>\n>I assert based on one piece of evidence ( a post from a facebook dev) \n>that I now have the worlds biggest and slowest git\n> repository, and I am not a happy guy. I used to have the worlds \n> biggest CVS repository, but CVS can't handle multi-G\n> sized files. So I moved the repo to git, because we are using that for \n> our new projects.\n>\n> goal:\n> keep 150 G of files (mostly binary) from tiny sized to over 8G in a \n> version-control system.\n>\n> problem:\n> git is absurdly slow, think hours, on fast hardware.\n>\n> question:\n> any suggestions beyond these-\n> http://git-annex.branchable.com/\n> https://github.com/jedbrown/git-fat\n> https://github.com/schacon/git-media\n> http://code.google.com/p/boar/\n> subversion\n>\n> ?\n\nAt the moment some of the developers are looking to speed up some of the \ncode on very large repos, though I think they are looking at code repos, \nrather than large file repos. They were looking for large repos to test \nsome of the code upon ;-)\n\nI've copied the Git list should they want to make any suggestions.\n\n>\n>\n> Thanks.\n>\n> -- \nPhilip \n\n-- \nYou received this message because you are subscribed to the Google Groups \"Git for human beings\" group.\nTo unsubscribe from this group and stop receiving emails from it, send an email to git-users+unsubscribe-/JYPxA39Uh5TLH3MbocFF+G/Ez6ZCGd0@public.gmane.org\nFor more options, visit https://groups.google.com/d/optout.\n"},{"id":"241779","messageId":"53751A0D.2020702@vilain.net","threadId":"36670","inReplyTo":"06A2490FC9BC4461A39B982D3C7C85F7@PhilipOakley","subject":"Re: [git-users] worlds slowest git repo- what to do?","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2014-05-15T19:48:29Z","receivedAt":"2014-05-15T19:48:29Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 05/15/2014 12:06 PM, Philip Oakley wrote:\n> From: \"John Fisher\" <fishook2033@gmail.com>\n>> I assert based on one piece of evidence ( a post from a facebook dev)\n>> that I now have the worlds biggest and slowest git\n>> repository, and I am not a happy guy. I used to have the worlds\n>> biggest CVS repository, but CVS can't handle multi-G\n>> sized files. So I moved the repo to git, because we are using that\n>> for our new projects.\n>>\n>> goal:\n>> keep 150 G of files (mostly binary) from tiny sized to over 8G in a\n>> version-control system.\n>>\n>> problem:\n>> git is absurdly slow, think hours, on fast hardware.\n>>\n>> question:\n>> any suggestions beyond these-\n>> http://git-annex.branchable.com/\n>> https://github.com/jedbrown/git-fat\n>> https://github.com/schacon/git-media\n>> http://code.google.com/p/boar/\n>> subversion\n>>\n\nYou could shard.  Break the problem up into smaller repositories, eg via\nsubmodules.  Try ~128 shards and I'd expect that 129 small clones should\ncomplete faster than a single 150G clone, as well as being resumable etc.\n\nThe first challenge will be figuring out what to shard on, and how to\nlay out the repository.  You could have all of the large files in their\nown directory, and then the main repository just has symlinks into the\nsharded area.  In that case, I would recommend sharding by date of the\nintroduced blob, so that there's a good chance you won't need to clone\neverything forever; as shards with not many files for the current\nversion could in theory be retired.  Or, if the directory structure\nalready suits it, you could \"directly\" use submodules.\n\nThe second challenge will be writing the filter-branch script for this :-)\n\nGood luck,\nSam\n"},{"id":"241872","messageId":"CACsJy8CmiW88tNavRphZa_uMU=jVUCQE6cw5+t2AYnf5dDmcsQ@mail.gmail.com","threadId":"36670","inReplyTo":"06A2490FC9BC4461A39B982D3C7C85F7@PhilipOakley","subject":"Re: [git-users] worlds slowest git repo- what to do?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2014-05-16T10:13:11Z","receivedAt":"2014-05-16T10:13:11Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, May 16, 2014 at 2:06 AM, Philip Oakley <philipoakley@iee.org> wrote:\n> From: \"John Fisher\" <fishook2033@gmail.com>\n>>\n>> I assert based on one piece of evidence ( a post from a facebook dev) that\n>> I now have the worlds biggest and slowest git\n>> repository, and I am not a happy guy. I used to have the worlds biggest\n>> CVS repository, but CVS can't handle multi-G\n>> sized files. So I moved the repo to git, because we are using that for our\n>> new projects.\n>>\n>> goal:\n>> keep 150 G of files (mostly binary) from tiny sized to over 8G in a\n>> version-control system.\n\nI think your best bet so far is git-annex (or maybe bup) for dealing\nwith huge files. I plan on resurrecting Junio's split-blob series to\nmake core git handle huge files better, but there's no eta on that.\nThe problem here is about file size, not the number of files, or\nhistory depth, right?\n\n>> problem:\n>> git is absurdly slow, think hours, on fast hardware.\n\nProbably known issues. But some elaboration would be nice (e.g. what\noperation is slow, how slow, some more detail characteristics of the\nrepo..) in case new problems pop up.\n-- \nDuy\n"},{"id":"242000","messageId":"537681A4.2070601@gmail.com","threadId":"36670","inReplyTo":"CACsJy8CmiW88tNavRphZa_uMU=jVUCQE6cw5+t2AYnf5dDmcsQ-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org","subject":"Re: worlds slowest git repo- what to do?","fromName":"John Fisher","fromEmail":"fishook2033-re5jqeeqqe8avxtiumwx3w@public.gmane.org","sentAt":"2014-05-16T21:22:44Z","receivedAt":"2014-05-16T21:22:44Z","isPatch":false,"sender":{"key":"fishook2033-re5jqeeqqe8avxtiumwx3w@public.gmane.org","avatar":null},"body":"\nOn 05/16/2014 03:13 AM, Duy Nguyen wrote:\n> On Fri, May 16, 2014 at 2:06 AM, Philip Oakley <philipoakley-7KbaBNvhQFM@public.gmane.org> wrote:\n>> From: \"John Fisher\" <fishook2033-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>\n>>> I assert based on one piece of evidence ( a post from a facebook dev) that\n>>> I now have the worlds biggest and slowest git\n>>> repository, and I am not a happy guy. I used to have the worlds biggest\n>>> CVS repository, but CVS can't handle multi-G\n>>> sized files. So I moved the repo to git, because we are using that for our\n>>> new projects.\n>>>\n>>> goal:\n>>> keep 150 G of files (mostly binary) from tiny sized to over 8G in a\n>>> version-control system.\n> I think your best bet so far is git-annex \n\ngood, I am  looking at that\n\n> (or maybe bup) for dealing\n> with huge files. I plan on resurrecting Junio's split-blob series to\n> make core git handle huge files better, but there's no eta on that.\n> The problem here is about file size, not the number of files, or\n> history depth, right?\n\nWhen things here calm down, I could easily test the repo without the giant files, leaving 99% of files in the repo.\nThere is hardly any history depth because these are releases, version controlled by directory name. As has been\nsuggested I could be forced to abandon the version-control, even to the point of just using rsync.  But I've been doing\nthis with CVS for 10 years now and I hate to change or in any way move away fron KISS. Moving it to Git may not have\nbeen one of my better ideas...\n\n\n> Probably known issues. But some elaboration would be nice (e.g. what operation is slow, how slow, some more detail\n> characteristics of the repo..) in case new problems pop up. \n\nso far I have done add, commit, status, clone - commit and status are slow; add seems to depend on the files involved,\nclone seems to run at network speed.\nI can provide metrics later, see above. email me offline with what you want.\n\nJohn\n\n-- \nYou received this message because you are subscribed to the Google Groups \"Git for human beings\" group.\nTo unsubscribe from this group and stop receiving emails from it, send an email to git-users+unsubscribe-/JYPxA39Uh5TLH3MbocFF+G/Ez6ZCGd0@public.gmane.org\nFor more options, visit https://groups.google.com/d/optout.\n"},{"id":"242016","messageId":"CACsJy8CmyJbUGAfb7SaAo5aNwOrJ8iwUtxZgQTAANChOsaWCRA@mail.gmail.com","threadId":"36670","inReplyTo":"537681A4.2070601@gmail.com","subject":"Re: [git-users] worlds slowest git repo- what to do?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2014-05-17T01:49:32Z","receivedAt":"2014-05-17T01:49:32Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, May 17, 2014 at 4:22 AM, John Fisher <fishook2033@gmail.com> wrote:\n>> Probably known issues. But some elaboration would be nice (e.g. what operation is slow, how slow, some more detail\n>> characteristics of the repo..) in case new problems pop up.\n>\n> so far I have done add, commit, status, clone - commit and status are slow; add seems to depend on the files involved,\n> clone seems to run at network speed.\n> I can provide metrics later, see above. email me offline with what you want.\n\nOK \"commit -a\" should be just as slow as \"add\", but as-is commit and\nstatus should be fast unless there are lots of files (how many in your\nworktree?) or we hit something that makes us look into (large) file\ncontent anyway.\n-- \nDuy\n"}]}