{"thread":{"id":"29536","subject":"Git performance results on a large repository","startedAt":"2012-02-03T14:20:06Z","lastAt":"2012-02-10T12:24:55Z","messageCount":34,"participants":["Joshua Redstone","Ævar Arnfjörð Bjarmason","Sam Vilain","Matt Graham","Chris Lee","Zeki Mokhtarzada","Evgeny Sazhin","Joey Hess","Nguyen Thai Ngoc Duy","slinky","Greg Troxel","david@lang.hm","David Barr","Tomas Carnecky","David Mohs","Emanuele Zattin","Christian Couder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"183734","messageId":"CB5074CF.3AD7A%joshua.redstone@fb.com","threadId":"29536","inReplyTo":null,"subject":"Git performance results on a large repository","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2012-02-03T14:20:06Z","receivedAt":"2012-02-03T14:20:06Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi Git folks,\n\nWe (Facebook) have been investigating source control systems to meet our\ngrowing needs.  We already use git fairly widely, but have noticed it\ngetting slower as we grow, and we want to make sure we have a good story\ngoing forward.  We're debating how to proceed and would like to solicit\npeople's thoughts.\n\nTo better understand git scalability, I've built up a large, synthetic\nrepository and measured a few git operations on it.  I summarize the\nresults here.\n\nThe test repo has 4 million commits, linear history and about 1.3 million\nfiles.  The size of the .git directory is about 15GB, and has been\nrepacked with 'git repack -a -d -f --max-pack-size=10g --depth=100\n--window=250'.  This repack took about 2 days on a beefy machine (I.e.,\nlots of ram and flash).  The size of the index file is 191 MB. I can share\nthe script that generated it if people are interested - It basically picks\n2-5 files, modifies a line or two and adds a few lines at the end\nconsisting of random dictionary words, occasionally creates a new file,\ncommits all the modifications and repeats.\n\nI timed a few common operations with both a warm OS file cache and a cold\ncache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did\nthe operation in question a few times (first timing is the cold timing,\nthe next few are the warm timings).  The following results are on a server\nwith average hard drive (I.e., not flash)  and > 10GB of ram.\n\n'git status' :   39 minutes cold, and 24 seconds warm.\n\n'git blame':   44 minutes cold, 11 minutes warm.\n\n'git add' (appending a few chars to the end of a file and adding it):   7\nseconds cold and 5 seconds warm.\n\n'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n--no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\nof git to remove the three or four places where 'git commit' stats every\nfile in the repo, and this dropped the times to 30 minutes cold and 8\nseconds warm.\n\n\nThe git performance we observed here is too slow for our needs.  So the\nquestion becomes, if we want to keep using git going forward, what's the\nbest way to improve performance.  It seems clear we'll probably need some\nspecialized servers (e.g., to perform git-blame quickly) and maybe\nspecialized file system integration to detect what files have changed in a\nworking tree.\n\nOne way to get there is to do some deep code modifications to git\ninternals, to, for example, create some abstractions and interfaces that\nallow plugging in the specialized servers.  Another way is to leave git\ninternals as they are and develop a layer of wrapper scripts around all\nthe git commands that do the necessary interfacing.  The wrapper scripts\nseem perhaps easier in the short-term, but may lead to increasing\ndivergence from how git behaves natively and also a layer of complexity.\n\nThoughts?\n\nCheers,\nJosh\n"},{"id":"183738","messageId":"CACBZZX4BsFZxB6A-Hg-k37FBavgTV8SDiQTK_sVh9Mb9iskiEw@mail.gmail.com","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2012-02-03T14:56:35Z","receivedAt":"2012-02-03T14:56:35Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Fri, Feb 3, 2012 at 15:20, Joshua Redstone <joshua.redstone@fb.com> wrote:\n\n> We (Facebook) have been investigating source control systems to meet our\n> growing needs.  We already use git fairly widely, but have noticed it\n> getting slower as we grow, and we want to make sure we have a good story\n> going forward.  We're debating how to proceed and would like to solicit\n> people's thoughts.\n\nWhere I work we also have a relatively large Git repository. Around\n30k files, a couple of hundred thousand commits, clone size around\nhalf a GB.\n\nYou haven't supplied background info on this but it really seems to me\nlike your testcase is converting something like a humongous Perforce\nrepository directly to Git.\n\nWhile you /can/ do this it's not a good idea, you should split up\nrepositories at the boundaries code or data doesn't directly cross\nover, e.g. there's no reason why you need HipHop PHP in the same\nrepository as Cassandra or the Facebook chat system, is there?\n\nWhile Git could better with large repositories (in particular applying\ncommits in interactive rebase seems to be to slow down on bigger\nrepositories) there's only so much you can do about stat-ing 1.3\nmillion files.\n\nA structure that would make more sense would be to split up that giant\nrepository into a lot of other repositories, most of them probably\nhave no direct dependencies on other components, but even those that\ndo can sometimes just use some other repository as a submodule.\n\nEven if you have the requirement that you'd like to roll out\n*everything* at a certain point in time you can still solve that with\na super-repository that has all the other ones as submodules, and\ncreates a tag for every rollout or something like that.\n"},{"id":"183742","messageId":"CB5179E9.3B751%joshua.redstone@fb.com","threadId":"29536","inReplyTo":"CACBZZX4BsFZxB6A-Hg-k37FBavgTV8SDiQTK_sVh9Mb9iskiEw@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2012-02-03T17:00:02Z","receivedAt":"2012-02-03T17:00:02Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi Ævar,\n\n\nThanks for the comments.  I've included a bunch more info on the test repo\nbelow.  It is based on a growth model of two of our current repositories\n(I.e., it's not a perforce import). We already have some of the easily\nseparable projects in separate repositories, like HPHP.   If we could\nsplit our largest repos into multiple ones, that would help the scaling\nissue.  However, the code in those repos is rather interdependent and we\nbelieve it'd hurt more than help to split it up, at least for the\nmedium-term future.  We derive a fair amount of benefit from the code\nsharing and keeping things together in a single repo, so it's not clear\nwhen it'd make sense to get more aggressive splitting things up.\n\nSome more information on the test repository:   The working directory is\n9.5 GB, the median file size is 2 KB.  The average depth of a directory\n(counting the number of '/'s) is 3.6 levels and the average depth of a\nfile is 4.6.  More detailed histograms of the repository composition is\nbelow:\n\n------------------------\n\nHistogram of depth of every directory in the repo (dirs=`find . -type d` ;\n(for dir in $dirs; do t=${dir//[^\\/]/}; echo ${#t} ; done) |\n~/tmp/histo.py)\n* The .git directory itself has only 161 files, so although included,\ndoesn't affect the numbers significantly)\n\n[0.0 - 1.3): 271\n[1.3 - 2.6): 9966\n[2.6 - 3.9): 56595\n[3.9 - 5.2): 230239\n[5.2 - 6.5): 67394\n[6.5 - 7.8): 22868\n[7.8 - 9.1): 6568\n[9.1 - 10.4): 420\n[10.4 - 11.7): 45\n[11.7 - 13.0]: 21\nn=394387 mean=4.671830, median=5.000000, stddev=1.272658\n\n\nHistogram of depth of every file in the repo (files=`git ls-files` ; (for\nfile in $files; do t=${file//[^\\/]/}; echo ${#t} ; done) | ~/tmp/histo.py)\n* 'git ls-files' does not prefix entries with ./, like the 'find' command\nabove, does, hence why the average appears to be the same as the directory\nstats\n\n[0.0 - 1.3]: 1274\n[1.3 - 2.6]: 35353\n[2.6 - 3.9]: 196747\n[3.9 - 5.2]: 786647\n[5.2 - 6.5]: 225913\n[6.5 - 7.8]: 77667\n[7.8 - 9.1]: 22130\n[9.1 - 10.4]: 1599\n[10.4 - 11.7]: 164\n[11.7 - 13.0]: 118\nn=1347612 mean=4.655750, median=5.000000, stddev=1.278399\n\n\nHistogram of file sizes (for first 50k files - this command takes a\nwhile):  files=`git ls-files` ; (for file in $files; do stat -c%s $file ;\ndone) | ~/tmp/histo.py\n\n[ 0.0 - 4.7): 0\n[ 4.7 - 22.5): 2\n[ 22.5 - 106.8): 0\n[ 106.8 - 506.8): 0\n[ 506.8 - 2404.7): 31142\n[ 2404.7 - 11409.9): 17837\n[ 11409.9 - 54137.1): 942\n[ 54137.1 - 256866.9): 53\n[ 256866.9 - 1218769.7): 18\n[ 1218769.7 - 5782760.0]: 5\nn=49999 mean=3590.953239, median=1772.000000, stddev=42835.330259\n\nCheers,\nJosh\n\n\n\n\n\n\nOn 2/3/12 9:56 AM, \"Ævar Arnfjörð Bjarmason\" <avarab@gmail.com> wrote:\n\n>On Fri, Feb 3, 2012 at 15:20, Joshua Redstone <joshua.redstone@fb.com>\n>wrote:\n>\n>> We (Facebook) have been investigating source control systems to meet our\n>> growing needs.  We already use git fairly widely, but have noticed it\n>> getting slower as we grow, and we want to make sure we have a good story\n>> going forward.  We're debating how to proceed and would like to solicit\n>> people's thoughts.\n>\n>Where I work we also have a relatively large Git repository. Around\n>30k files, a couple of hundred thousand commits, clone size around\n>half a GB.\n>\n>You haven't supplied background info on this but it really seems to me\n>like your testcase is converting something like a humongous Perforce\n>repository directly to Git.\n>\n>While you /can/ do this it's not a good idea, you should split up\n>repositories at the boundaries code or data doesn't directly cross\n>over, e.g. there's no reason why you need HipHop PHP in the same\n>repository as Cassandra or the Facebook chat system, is there?\n>\n>While Git could better with large repositories (in particular applying\n>commits in interactive rebase seems to be to slow down on bigger\n>repositories) there's only so much you can do about stat-ing 1.3\n>million files.\n>\n>A structure that would make more sense would be to split up that giant\n>repository into a lot of other repositories, most of them probably\n>have no direct dependencies on other components, but even those that\n>do can sometimes just use some other repository as a submodule.\n>\n>Even if you have the requirement that you'd like to roll out\n>*everything* at a certain point in time you can still solve that with\n>a super-repository that has all the other ones as submodules, and\n>creates a tag for every rollout or something like that.\n\n"},{"id":"183776","messageId":"4F2C6276.1070100@vilain.net","threadId":"29536","inReplyTo":"CB5179E9.3B751%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2012-02-03T22:40:54Z","receivedAt":"2012-02-03T22:40:54Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Joshua,\n\nYou have an interesting use case.\n\nIf I were you I'd consider investigating the git fast-import protocol. \nIt has become bi–directional, and is essentially socket access to a git \nrepository with read and transactional update capability.  With a few \nmore commands implemented, it may even be capable of providing all \nfunctionality required for command–line git use.\n\nIt is already possible that the \".git\" directory can be a file: this \ncase is used for submodules in git 1.7.8 and higher.  For this use case, \nthere would be an extra field to the \".git\" file which is created.  It \nwould indicate the hostname (and port) to connect its internal \n'fast-import' stream to.  'clone' would consist of creating this file, \nand then getting the server to stream the objects from its pack to the \nclient.\n\nWith the hard–working part of git on the other end of a network service, \nyou could back it by a re–implementation of git which is written to be \ndistributed in Hadoop.  There are at least two similar implementations \nof git that are like this: one for cassandra which was written by github \nas a research project, and Google's implementation on top of their \nBigTable/GFS/whatever.  As the git object storage model is write–only \nand content–addressed, it should git this kind of scaling well.\n\nThere have also been designs at various times for sparse check–outs; ie \ncheck–outs where you don't check out the root of the repository but a \nsub–tree.\n\nWith both of these features, clients could easily check out a small part \nof the repository very quickly.  This is probably the only case which \nSVN still does better than git at, which is a particular blocker for use \ncases like repositories with large binaries in them and for projects \nsuch as the one you have (another one with a similar problem was KDE, \nwhere their projects moved around the repository a lot, and refactoring \ntouched many projects simultaneously at times).\n\nIt's a large undertaking, alright.\n\nSam,\njust another git community propeller–head.\n\n\nOn 2/3/12 9:00 AM, Joshua Redstone wrote:\n> Hi Ævar,\n>\n>\n> Thanks for the comments.  I've included a bunch more info on the test repo\n> below.  It is based on a growth model of two of our current repositories\n> (I.e., it's not a perforce import). We already have some of the easily\n> separable projects in separate repositories, like HPHP.   If we could\n> split our largest repos into multiple ones, that would help the scaling\n> issue.  However, the code in those repos is rather interdependent and we\n> believe it'd hurt more than help to split it up, at least for the\n> medium-term future.  We derive a fair amount of benefit from the code\n> sharing and keeping things together in a single repo, so it's not clear\n> when it'd make sense to get more aggressive splitting things up.\n>\n> Some more information on the test repository:   The working directory is\n> 9.5 GB, the median file size is 2 KB.  The average depth of a directory\n> (counting the number of '/'s) is 3.6 levels and the average depth of a\n> file is 4.6.  More detailed histograms of the repository composition is\n> below:\n>\n> ------------------------\n>\n> Histogram of depth of every directory in the repo (dirs=`find . -type d` ;\n> (for dir in $dirs; do t=${dir//[^\\/]/}; echo ${#t} ; done) |\n> ~/tmp/histo.py)\n> * The .git directory itself has only 161 files, so although included,\n> doesn't affect the numbers significantly)\n>\n> [0.0 - 1.3): 271\n> [1.3 - 2.6): 9966\n> [2.6 - 3.9): 56595\n> [3.9 - 5.2): 230239\n> [5.2 - 6.5): 67394\n> [6.5 - 7.8): 22868\n> [7.8 - 9.1): 6568\n> [9.1 - 10.4): 420\n> [10.4 - 11.7): 45\n> [11.7 - 13.0]: 21\n> n=394387 mean=4.671830, median=5.000000, stddev=1.272658\n>\n>\n> Histogram of depth of every file in the repo (files=`git ls-files` ; (for\n> file in $files; do t=${file//[^\\/]/}; echo ${#t} ; done) | ~/tmp/histo.py)\n> * 'git ls-files' does not prefix entries with ./, like the 'find' command\n> above, does, hence why the average appears to be the same as the directory\n> stats\n>\n> [0.0 - 1.3]: 1274\n> [1.3 - 2.6]: 35353\n> [2.6 - 3.9]: 196747\n> [3.9 - 5.2]: 786647\n> [5.2 - 6.5]: 225913\n> [6.5 - 7.8]: 77667\n> [7.8 - 9.1]: 22130\n> [9.1 - 10.4]: 1599\n> [10.4 - 11.7]: 164\n> [11.7 - 13.0]: 118\n> n=1347612 mean=4.655750, median=5.000000, stddev=1.278399\n>\n>\n> Histogram of file sizes (for first 50k files - this command takes a\n> while):  files=`git ls-files` ; (for file in $files; do stat -c%s $file ;\n> done) | ~/tmp/histo.py\n>\n> [ 0.0 - 4.7): 0\n> [ 4.7 - 22.5): 2\n> [ 22.5 - 106.8): 0\n> [ 106.8 - 506.8): 0\n> [ 506.8 - 2404.7): 31142\n> [ 2404.7 - 11409.9): 17837\n> [ 11409.9 - 54137.1): 942\n> [ 54137.1 - 256866.9): 53\n> [ 256866.9 - 1218769.7): 18\n> [ 1218769.7 - 5782760.0]: 5\n> n=49999 mean=3590.953239, median=1772.000000, stddev=42835.330259\n>\n> Cheers,\n> Josh\n>\n>\n>\n>\n>\n>\n> On 2/3/12 9:56 AM, \"Ævar Arnfjörð Bjarmason\"<avarab@gmail.com>  wrote:\n>\n>> On Fri, Feb 3, 2012 at 15:20, Joshua Redstone<joshua.redstone@fb.com>\n>> wrote:\n>>\n>>> We (Facebook) have been investigating source control systems to meet our\n>>> growing needs.  We already use git fairly widely, but have noticed it\n>>> getting slower as we grow, and we want to make sure we have a good story\n>>> going forward.  We're debating how to proceed and would like to solicit\n>>> people's thoughts.\n>>\n>> Where I work we also have a relatively large Git repository. Around\n>> 30k files, a couple of hundred thousand commits, clone size around\n>> half a GB.\n>>\n>> You haven't supplied background info on this but it really seems to me\n>> like your testcase is converting something like a humongous Perforce\n>> repository directly to Git.\n>>\n>> While you /can/ do this it's not a good idea, you should split up\n>> repositories at the boundaries code or data doesn't directly cross\n>> over, e.g. there's no reason why you need HipHop PHP in the same\n>> repository as Cassandra or the Facebook chat system, is there?\n>>\n>> While Git could better with large repositories (in particular applying\n>> commits in interactive rebase seems to be to slow down on bigger\n>> repositories) there's only so much you can do about stat-ing 1.3\n>> million files.\n>>\n>> A structure that would make more sense would be to split up that giant\n>> repository into a lot of other repositories, most of them probably\n>> have no direct dependencies on other components, but even those that\n>> do can sometimes just use some other repository as a submodule.\n>>\n>> Even if you have the requirement that you'd like to roll out\n>> *everything* at a certain point in time you can still solve that with\n>> a super-repository that has all the other ones as submodules, and\n>> creates a tag for every rollout or something like that.\n>\n> N�����r��y���b�X��ǧv�^�)޺{.n�+����ا�\u0017��ܨ}���Ơz�&j:+v���\u0007����zZ+��+zf���h���~����i���z�\u001e�w���?����&�)ߢ\u001bfl===\n"},{"id":"183778","messageId":"4F2C665C.8080909@vilain.net","threadId":"29536","inReplyTo":"4F2C6276.1070100@vilain.net","subject":"Re: Git performance results on a large repository","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2012-02-03T22:57:32Z","receivedAt":"2012-02-03T22:57:32Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"On 2/3/12 2:40 PM, Sam Vilain wrote:\n> As the git object storage model is write–only and content–addressed,\n> it should git this kind of scaling well.\n             ^^^\n\nCould have sworn I typed 'suit' there.  My fingers have auto–correct ;-)\n\nSam\n"},{"id":"183782","messageId":"CALts4TT49VAWPZ6XO9qahDTu=2E425QcRvXx5-75Jv8n4yp8RA@mail.gmail.com","threadId":"29536","inReplyTo":"CB5179E9.3B751%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Matt Graham","fromEmail":"mdg149@gmail.com","sentAt":"2012-02-03T23:05:40Z","receivedAt":"2012-02-03T23:05:40Z","isPatch":false,"sender":{"key":"mdg149@gmail.com","avatar":"https://gravatar.com/avatar/a1f130a60a6550f75e8d7d3849e58e46494f36bfaf764a38cfd695ac85de8576?d=mp&s=160"},"body":"Hi Josh,\n\nOn Fri, Feb 3, 2012 at 17:00, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> Thanks for the comments.  I've included a bunch more info on the test repo\n> below.  It is based on a growth model of two of our current repositories\n> (I.e., it's not a perforce import). We already have some of the easily\n> separable projects in separate repositories, like HPHP.   If we could\n> split our largest repos into multiple ones, that would help the scaling\n> issue.  However, the code in those repos is rather interdependent and we\n> believe it'd hurt more than help to split it up, at least for the\n> medium-term future.  We derive a fair amount of benefit from the code\n> sharing and keeping things together in a single repo, so it's not clear\n> when it'd make sense to get more aggressive splitting things up.\n>\n> Some more information on the test repository:   The working directory is\n> 9.5 GB, the median file size is 2 KB.  The average depth of a directory\n> (counting the number of '/'s) is 3.6 levels and the average depth of a\n> file is 4.6.  More detailed histograms of the repository composition is\n> below:\n\nDo you have a histogram of the types of files in the repo?\nAnd as suggested earlier, is svn working for you now because it allows\nsparse checkout?  I imagine the stats for svn on the full repo would\nbe comparable or worse to what you measured with git?\n"},{"id":"183789","messageId":"CA+B68xBG8c3eMg5ULUYgmZ4vTiSgvu2nEFfgTD1m0-dLhLKZhg@mail.gmail.com","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2012-02-03T23:35:05Z","receivedAt":"2012-02-03T23:35:05Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On Fri, Feb 3, 2012 at 6:20 AM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> [snip]\n>\n> The git performance we observed here is too slow for our needs.  So the\n> question becomes, if we want to keep using git going forward, what's the\n> best way to improve performance.  It seems clear we'll probably need some\n> specialized servers (e.g., to perform git-blame quickly) and maybe\n> specialized file system integration to detect what files have changed in a\n> working tree.\n\nHave you considered upgrading all of engineering to SSDs? 200+GB SSDs\nare under $400USD nowadays.\n\n-clee\n"},{"id":"183791","messageId":"loom.20120204T004543-798@post.gmane.org","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Zeki Mokhtarzada","fromEmail":"zeki@webs.com","sentAt":"2012-02-04T00:01:18Z","receivedAt":"2012-02-04T00:01:18Z","isPatch":false,"sender":{"key":"zeki@webs.com","avatar":null},"body":" \n> The test repo has 4 million commits, linear history and about 1.3 million\n> files.  The size of the .git directory is about 15GB, and has been\n> repacked with 'git repack -a -d -f --max-pack-size=10g --depth=100\n> --window=250'.  This repack took about 2 days on a beefy machine (I.e.,\n> lots of ram and flash).  The size of the index file is 191 MB. I can share\n\n\nAre you willing to give up all or part of your history in your working\nrepository?  I've heard of larger projects starting from scratch (i.e. copy all\nof your files into a brand new repo.)  You can keep your old repo around for\narchival purposes.  Also, how much of your repo is code, versus static assets. \nYou could move all of your static assets (images, css, maybe some js?) into\nanother repo, and then merge the two repo's together at build time if you\nabsolutely need them deployed together.\n\nHere are a couple strategies for doing a partial truncate:\n\nhttp://stackoverflow.com/questions/4515580/how-do-i-remove-the-old-history-from-a-git-repository\nhttp://bogdan.org.ua/2011/03/28/how-to-truncate-git-history-sample-script-included.html\n\n\n-Zeki\n"},{"id":"183793","messageId":"6E708713-3DEF-40A2-9585-690707166BDF@gmail.com","threadId":"29536","inReplyTo":"CACBZZX4BsFZxB6A-Hg-k37FBavgTV8SDiQTK_sVh9Mb9iskiEw@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Evgeny Sazhin","fromEmail":"euguess@gmail.com","sentAt":"2012-02-04T01:25:59Z","receivedAt":"2012-02-04T01:25:59Z","isPatch":false,"sender":{"key":"euguess@gmail.com","avatar":null},"body":" \n\nOn Feb 3, 2012, at 9:56 AM, Ævar Arnfjörð Bjarmason wrote:\n\n> On Fri, Feb 3, 2012 at 15:20, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> \n>> We (Facebook) have been investigating source control systems to meet our\n>> growing needs.  We already use git fairly widely, but have noticed it\n>> getting slower as we grow, and we want to make sure we have a good story\n>> going forward.  We're debating how to proceed and would like to solicit\n>> people's thoughts.\n> \n> Where I work we also have a relatively large Git repository. Around\n> 30k files, a couple of hundred thousand commits, clone size around\n> half a GB.\n> \n> You haven't supplied background info on this but it really seems to me\n> like your testcase is converting something like a humongous Perforce\n> repository directly to Git.\n> \n> While you /can/ do this it's not a good idea, you should split up\n> repositories at the boundaries code or data doesn't directly cross\n> over, e.g. there's no reason why you need HipHop PHP in the same\n> repository as Cassandra or the Facebook chat system, is there?\n> \n> While Git could better with large repositories (in particular applying\n> commits in interactive rebase seems to be to slow down on bigger\n> repositories) there's only so much you can do about stat-ing 1.3\n> million files.\n> \n> A structure that would make more sense would be to split up that giant\n> repository into a lot of other repositories, most of them probably\n> have no direct dependencies on other components, but even those that\n> do can sometimes just use some other repository as a submodule.\n> \n> Even if you have the requirement that you'd like to roll out\n> *everything* at a certain point in time you can still solve that with\n> a super-repository that has all the other ones as submodules, and\n> creates a tag for every rollout or something like that.\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n\n\nI concur. I'm working in the company with many years of development history with several huge CVS repos and we are slowly but surely migrating the codebase from CVS to Git. \nSplit the things up. This will allow you to reorganize things better and there is IMHO no downsides. \nAs for rollout - i think this job should be given to build/release system that will have an ability to gather necessary code from different repos and tag it properly.\n\njust my 2 cents\n\nThanks,\nEugene\n"},{"id":"183798","messageId":"20120204050712.GA2460@gnu.kitenet.net","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Joey Hess","fromEmail":"joey@kitenet.net","sentAt":"2012-02-04T05:07:12Z","receivedAt":"2012-02-04T05:07:12Z","isPatch":false,"sender":{"key":"joey@kitenet.net","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Joshua Redstone wrote:\n> The test repo has 4 million commits, linear history and about 1.3 million\n> files.\n\nHave you tried separating these two factors, to see how badly each is\naffecting performance?\n\nIf the number of commits is the problem (seems likely for git blame at\nleast), a shallow clone would avoid that overhead.\n\nI think that git often writes .git/index inneficiently when staging\nfiles (though your `git add` is pretty fast) and committing. It rewrites\nthe whole file to .git/index.lck and the renames it over .git/index at\nthe end. I have code that keeps a journal of changes to avoid rewriting\nthe index repeatedly, but it's application specific. Fixing git to write\nthe index more intelligently is something I'd like to see.\n\nHint for git status: `git status .` in a smaller subdirectory will be much\nfaster than the default that stats everything.\n\n-- \nsee shy jo\n"},{"id":"183810","messageId":"CACsJy8DkLCK0ZUKNz_PJazsxjsRbWVVZwjAU5n2EAjJfCYtpoQ@mail.gmail.com","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-04T06:53:47Z","receivedAt":"2012-02-04T06:53:47Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> I timed a few common operations with both a warm OS file cache and a cold\n> cache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did\n> the operation in question a few times (first timing is the cold timing,\n> the next few are the warm timings).  The following results are on a server\n> with average hard drive (I.e., not flash)  and > 10GB of ram.\n>\n> 'git status' :   39 minutes cold, and 24 seconds warm.\n>\n> 'git blame':   44 minutes cold, 11 minutes warm.\n>\n> 'git add' (appending a few chars to the end of a file and adding it):   7\n> seconds cold and 5 seconds warm.\n>\n> 'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n> --no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\n> of git to remove the three or four places where 'git commit' stats every\n> file in the repo, and this dropped the times to 30 minutes cold and 8\n> seconds warm.\n\nHave you tried \"git update-index --assume-unchaged\"? That should\nreduce mass lstat() and hopefully improve the above numbers. The\ninterface is not exactly easy-to-use, but if it has significant gain,\nthen we can try to improve UI.\n\nOn the index size issue, ideally we should make minimum writes to\nindex instead of rewriting 191 MB index. An improvement we could do\nnow is to compress it, reduce disk footprint, thus disk I/O. If you\ncompress the index with gzip, how big is it?\n-- \nDuy\n"},{"id":"183818","messageId":"loom.20120204T094226-583@post.gmane.org","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"slinky","fromEmail":"slinky@iki.fi","sentAt":"2012-02-04T08:57:33Z","receivedAt":"2012-02-04T08:57:33Z","isPatch":false,"sender":{"key":"slinky@iki.fi","avatar":null},"body":"Joshua Redstone <joshua.redstone <at> fb.com> writes:\n\n> The git performance we observed here is too slow for our needs.  So the\n> question becomes, if we want to keep using git going forward, what's the\n> best way to improve performance.  It seems clear we'll probably need some\n> specialized servers (e.g., to perform git-blame quickly) and maybe\n> specialized file system integration to detect what files have changed in a\n> working tree.\n\nHi Joshua,\n\nsounds like you have everything in a single .git. Split up the massive\nrepository to separate smaller .git repositories.\n\nFor example, Android code base is quite big. They use the repo tool to manage a\nnumber of separate .git repositories as one big aggregate \"repository\".\n\nCheers,\nSlinky\n"},{"id":"183847","messageId":"243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com","threadId":"29536","inReplyTo":"CACsJy8DkLCK0ZUKNz_PJazsxjsRbWVVZwjAU5n2EAjJfCYtpoQ@mail.gmail.com","subject":"RE: Git performance results on a large repository","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2012-02-04T18:05:08Z","receivedAt":"2012-02-04T18:05:08Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"[ wanted to reply to my initial msg, but wasn't subscribed to the list at time of mailing, so replying to most recent post instead ]\n\nThanks to everyone for the questions and suggestions.  I'll try to respond here.  One high-level clarification - this synthetic repo for which I've reported perf times is representative of where we think we'll be in the future.  Git is slow but marginally acceptable for today.  We want to start planning now for any big changes we need to make going forward.\n\nEvgeny Sazhin, Slinky and Ævar Arnfjörð Bjarmason suggested splitting up the repo into multiple, smaller repos.  I indicated before that we have a lot of cross-dependencies.  Our largest repo by number of files and commits is the repo containing the front-end server.  It is a large code base in which the tight integration of various components results in many of the cross dependencies.  We are working slowly to split things up more, for example into services, but that is a long-term process.\n\nTo get a bit abstract for a moment, in an ideal world, it doesn't seem like performance constraints of a source-control-system should dictate how we choose to structure our code.  Ideally, seems like we should be able to choose to structure our code in whatever way we feel maximizes developer productivity.  If development and code/release management seem easier in a single repo, than why not make an SCM that can handle it?  This is one reason I've been leaning towards figuring out an SCM approach that can work well with our current practices rather than changing them as a prerequisite for good SCM performance.\n\nSam Vilain:  Thanks for the pointer, i didn't realize that fast-import was bi-directional.  I used it for generating the synthetic repo.  Will look into using it the other way around.  Though that still won't speed up things like git-blame, presumably?  The sparse-checkout issue you mention is a good one.  There is a good question of how to support quick checkout, branch switching, clone, push and so forth.  I'll look into the approaches you suggest.  One consideration is coming up with a high-leverage approach - i.e. not doing heavy dev work if we can avoid it.  On the other hand, it would be nice if we (including the entire community :) ) improve git in areas that others that share similar issues benefit from as well.\n\nMatt Graham:  I don't have file stats at the moment.  It's mostly code files, with a few larger data files here and there.    We also don't do sparse checkouts, primarily because most people use git (whether on top of SVN or not), which doesn't support it.\n\nChris Lee:  When I was building up the repo (e.g., doing lots of commits, before I started using fast-import), i noticed that flash was not much faster - stat'ing the whole repo takes a lot of kernel time, even with flash.  My hunch is that we'd see similar issues with other operations, like git-blame.\n\nZeki Mokhtarzada:  Dumping history I think would speed up operations for which we don't care about old history, like git-blame in which we only want to see recent modifications.  We'd also need a good story for other kinds of operations.  In my mental model of git scalability, I categorize git structures into three kinds:  those for reasoning about history, those for the index and those for the working directory  (yeah, I know these don't map precisely to actual on-disk things like the object store, including trees, etc.).  One scaling approach we've been thinking of is to focus on each individually:  develop a specialized thing to handle history commands efficiently (git-blame, git-log, git-diff, etc.), something to speed up or bypass the index, and something to make large changes to the working directly quickly.\n\nJoey Hess:  Separating the factors is a good suggestion.  My hunch is that the various git operations test the performance issues in isolation.  For example, git-status performance depends just on the number of files, not on the depth of history.  On the other hand, my guess is that git-blame performance is more a function of the length of history rather than the number of files.  Though, certainly with compression and indexing in pack files, I could imagine there being cross-effects between length of history and number of files.   The git-status suggestion definitely helps when you know which directory you are concerned about.  Often I'm lazy and stat the repo root so I trade-off slowness for being more sure I'm not missing anything.\n\n@Joey, I think you're also touching on a good meta point which is that, there's probably no silver bullet here.  If we want git to efficiently handle repos that are large across a number of dimensions (size, # commits, # files, etc.), there's multiple parts of git that would need enhancement of some form.\n\nNguyen Thai Ngoc Duy:  At which point in the test flow should I insert git-update-index?  I'm happy to try it out.  Will compress index when I next get to a terminal.  My guess is it'll compress a bunch.  It's also conceivable that, if there were an external interface in git to attach other systems to efficiently report which files have changed (e.g., via file-system integration), it's possible that we could omit managing the index in many cases.   I know that would be a big change, but the benefits are intriguing.\n\nCheers,\nJosh\n\n\n\n\n________________________________________\nFrom: Nguyen Thai Ngoc Duy [pclouds@gmail.com]\nSent: Friday, February 03, 2012 10:53 PM\nTo: Joshua Redstone\nCc: git@vger.kernel.org\nSubject: Re: Git performance results on a large repository\n\nOn Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> I timed a few common operations with both a warm OS file cache and a cold\n> cache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did\n> the operation in question a few times (first timing is the cold timing,\n> the next few are the warm timings).  The following results are on a server\n> with average hard drive (I.e., not flash)  and > 10GB of ram.\n>\n> 'git status' :   39 minutes cold, and 24 seconds warm.\n>\n> 'git blame':   44 minutes cold, 11 minutes warm.\n>\n> 'git add' (appending a few chars to the end of a file and adding it):   7\n> seconds cold and 5 seconds warm.\n>\n> 'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n> --no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\n> of git to remove the three or four places where 'git commit' stats every\n> file in the repo, and this dropped the times to 30 minutes cold and 8\n> seconds warm.\n\nHave you tried \"git update-index --assume-unchaged\"? That should\nreduce mass lstat() and hopefully improve the above numbers. The\ninterface is not exactly easy-to-use, but if it has significant gain,\nthen we can try to improve UI.\n\nOn the index size issue, ideally we should make minimum writes to\nindex instead of rewriting 191 MB index. An improvement we could do\nnow is to compress it, reduce disk footprint, thus disk I/O. If you\ncompress the index with gzip, how big is it?\n--\nDuy\n"},{"id":"183857","messageId":"243C23AF01622E49BEA3F28617DBF0AD5912CC95@SC-MBX02-5.TheFacebook.com","threadId":"29536","inReplyTo":"CACsJy8DkLCK0ZUKNz_PJazsxjsRbWVVZwjAU5n2EAjJfCYtpoQ@mail.gmail.com","subject":"RE: Git performance results on a large repository","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2012-02-04T20:05:11Z","receivedAt":"2012-02-04T20:05:11Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"One more follow-on thought.  I imagine that most consumers of git are nowhere near the scale of the test repo that I described.  They may still enjoy benefit from efforts to improve git support for large repos.  A few possible reasons:\n\n1. The performance improvements should speed things up for smaller repos as well.\n2. They may find their repos growing to a 'large scale' at some point in the future.\n3. Any code cleanup as part of an effort to support git scalability is good for code base health and e.g., would facilitate future modifications that may more directly affect them.\n\nCheers,\nJosh\n________________________________________\nFrom: Nguyen Thai Ngoc Duy [pclouds@gmail.com]\nSent: Friday, February 03, 2012 10:53 PM\nTo: Joshua Redstone\nCc: git@vger.kernel.org\nSubject: Re: Git performance results on a large repository\n\nOn Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> I timed a few common operations with both a warm OS file cache and a cold\n> cache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did\n> the operation in question a few times (first timing is the cold timing,\n> the next few are the warm timings).  The following results are on a server\n> with average hard drive (I.e., not flash)  and > 10GB of ram.\n>\n> 'git status' :   39 minutes cold, and 24 seconds warm.\n>\n> 'git blame':   44 minutes cold, 11 minutes warm.\n>\n> 'git add' (appending a few chars to the end of a file and adding it):   7\n> seconds cold and 5 seconds warm.\n>\n> 'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n> --no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\n> of git to remove the three or four places where 'git commit' stats every\n> file in the repo, and this dropped the times to 30 minutes cold and 8\n> seconds warm.\n\nHave you tried \"git update-index --assume-unchaged\"? That should\nreduce mass lstat() and hopefully improve the above numbers. The\ninterface is not exactly easy-to-use, but if it has significant gain,\nthen we can try to improve UI.\n\nOn the index size issue, ideally we should make minimum writes to\nindex instead of rewriting 191 MB index. An improvement we could do\nnow is to compress it, reduce disk footprint, thus disk I/O. If you\ncompress the index with gzip, how big is it?\n--\nDuy\n"},{"id":"183870","messageId":"rmivcnm2s3w.fsf@fnord.ir.bbn.com","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Greg Troxel","fromEmail":"gdt@ir.bbn.com","sentAt":"2012-02-04T21:42:11Z","receivedAt":"2012-02-04T21:42:11Z","isPatch":false,"sender":{"key":"gdt@ir.bbn.com","avatar":null},"body":"\nJoshua Redstone <joshua.redstone@fb.com> writes:\n\n> The test repo has 4 million commits, linear history and about 1.3 million\n> files.  The size of the .git directory is about 15GB, and has been\n> repacked with 'git repack -a -d -f --max-pack-size=10g --depth=100\n> --window=250'.  This repack took about 2 days on a beefy machine (I.e.,\n> lots of ram and flash).  The size of the index file is 191 MB. I can share\n> the script that generated it if people are interested - It basically picks\n> 2-5 files, modifies a line or two and adds a few lines at the end\n> consisting of random dictionary words, occasionally creates a new file,\n> commits all the modifications and repeats.\n\nI have a repository with about 500K files, 3.3G checkout, 1.5G .git, and\nabout 10K commits.  (This is a real repository, not a test case.)  So\nnot as many commits by a lot, but the size seems not so far off.\n\n> I timed a few common operations with both a warm OS file cache and a cold\n> cache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did\n> the operation in question a few times (first timing is the cold timing,\n> the next few are the warm timings).  The following results are on a server\n> with average hard drive (I.e., not flash)  and > 10GB of ram.\n>\n> 'git status' :   39 minutes cold, and 24 seconds warm.\n\nBoth of these numbers surprise me.  I'm using NetBSD, whose stat\nimplementation isn't as optimized as Linux (you didn't say, but\nassuming).   On a years-old desktop, git status seems to be about a\nminute semi-cold and 5s warm (once I set the vnode cache big over 500K,\nvs 350K default for a 2G ram machine).\n\nSo on the warm status, I wonder how big your vnode cache is, and if\nyou've exceeded it, and I don't follow the cold time at all.  Probably\nsome sort of profiling within git status would be illuminating.\n\n> 'git blame':   44 minutes cold, 11 minutes warm.\n>\n> 'git add' (appending a few chars to the end of a file and adding it):   7\n> seconds cold and 5 seconds warm.\n>\n> 'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n> --no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\n> of git to remove the three or four places where 'git commit' stats every\n> file in the repo, and this dropped the times to 30 minutes cold and 8\n> seconds warm.\n\nSo without the stat, I wonder what it's doing that takes 30 minutes.\n\n> One way to get there is to do some deep code modifications to git\n> internals, to, for example, create some abstractions and interfaces that\n> allow plugging in the specialized servers.  Another way is to leave git\n> internals as they are and develop a layer of wrapper scripts around all\n> the git commands that do the necessary interfacing.  The wrapper scripts\n> seem perhaps easier in the short-term, but may lead to increasing\n> divergence from how git behaves natively and also a layer of complexity.\n\nHaving hooks for a blame server cache, etc. sounds sensible.  Having a\nway to call blames sort of like with --since and then keep updating it\n(eg. in emacs) to earlier times sounds useful.\n"},{"id":"183884","messageId":"CACsJy8Bf95JMp1qOiruR7+Tdi7JN42KNeMqGLud+z3O26DREnw@mail.gmail.com","threadId":"29536","inReplyTo":"243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-05T03:47:27Z","receivedAt":"2012-02-05T03:47:27Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 5, 2012 at 1:05 AM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> It's also conceivable that, if there were an external interface in git to attach other\n> systems to efficiently report which files have changed (e.g., via file-system integration),\n> it's possible that we could omit managing the index in many cases.\n> I know that would be a big change, but the benefits are intriguing.\n\nThe \"interface to report which files have changed\" is exactly \"git\nupdate-index --[no-]assume-unchanged\" is for. Have a look at the man\npage. Basically you can mark every file \"unchanged\" in the beginning\nand git won't bother lstat() them. What files you change, you have to\nexplicitly run \"git update-index --no-assume-unchanged\" to tell git.\n\nSomeone on HN suggested making assume-unchanged files read-only to\navoid 90% accidentally changing a file without telling git. When\nassume-unchanged bit is cleared, the file is made read-write again.\n-- \nDuy\n"},{"id":"183885","messageId":"alpine.DEB.2.02.1202042026280.6541@asgard.lang.hm","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"","fromEmail":"david@lang.hm","sentAt":"2012-02-05T04:30:55Z","receivedAt":"2012-02-05T04:30:55Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 3 Feb 2012, Joshua Redstone wrote:\n\n> The test repo has 4 million commits, linear history and about 1.3 million\n> files.  The size of the .git directory is about 15GB, and has been\n> repacked with 'git repack -a -d -f --max-pack-size=10g --depth=100\n> --window=250'.  This repack took about 2 days on a beefy machine (I.e.,\n> lots of ram and flash).  The size of the index file is 191 MB.\n\nThis may be a silly thought, but what if instead of one pack file of your \nentire history (4 million commits) you create multiple packs (say every \nhalf million commits) and mark all but the most recent pack as .keep (so \nthat they won't be modified by a repack)\n\nthat way things that only need to worry about recent history (blame, etc) \nwill probably never have to go past the most recent pack file or two\n\nI may be wrong, but I think that when git is looking for 'similar files' \nfor delta compression, it limits it's search to the current pack, so this \nwill also keep you from searching the entire project history.\n\nDavid Lang\n"},{"id":"183895","messageId":"CAFfmPPMcS5Y8hoXyR8LY637z_f-A1XgrH6m+YUsCv8gmGjtr3w@mail.gmail.com","threadId":"29536","inReplyTo":"alpine.DEB.2.02.1202042026280.6541@asgard.lang.hm","subject":"Re: Git performance results on a large repository","fromName":"David Barr","fromEmail":"davidbarr@google.com","sentAt":"2012-02-05T11:24:45Z","receivedAt":"2012-02-05T11:24:45Z","isPatch":false,"sender":{"key":"davidbarr@google.com","avatar":"https://avatars.githubusercontent.com/u/220594?v=4"},"body":"On Sun, Feb 5, 2012 at 3:30 PM,  <david@lang.hm> wrote:\n> On Fri, 3 Feb 2012, Joshua Redstone wrote:\n>\n>> The test repo has 4 million commits, linear history and about 1.3 million\n>> files.  The size of the .git directory is about 15GB, and has been\n>> repacked with 'git repack -a -d -f --max-pack-size=10g --depth=100\n>> --window=250'.  This repack took about 2 days on a beefy machine (I.e.,\n>> lots of ram and flash).  The size of the index file is 191 MB.\n>\n>\n> This may be a silly thought, but what if instead of one pack file of your\n> entire history (4 million commits) you create multiple packs (say every half\n> million commits) and mark all but the most recent pack as .keep (so that\n> they won't be modified by a repack)\n>\n> that way things that only need to worry about recent history (blame, etc)\n> will probably never have to go past the most recent pack file or two\n>\n> I may be wrong, but I think that when git is looking for 'similar files' for\n> delta compression, it limits it's search to the current pack, so this will\n> also keep you from searching the entire project history.\n\nI don't know if there is an easy way to determine with the with the\ncurrent tools\nin git but one useful statistic for tuning packing performance is the\nsize of the\nlargest component in the delta-chain graph. The significance of this number is\nthat the product of window-size and maximum depth need not be larger than it.\nI've found that with some older repositories I could have a depth as low as 3\nand still get good performance from a moderate window size.\n\n--\nDavid Barr\n"},{"id":"183901","messageId":"4F2E99C2.7090609@dbservice.com","threadId":"29536","inReplyTo":"CACsJy8DkLCK0ZUKNz_PJazsxjsRbWVVZwjAU5n2EAjJfCYtpoQ@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2012-02-05T15:01:22Z","receivedAt":"2012-02-05T15:01:22Z","isPatch":false,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"On 2/4/12 7:53 AM, Nguyen Thai Ngoc Duy wrote:\n> On Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone<joshua.redstone@fb.com>  wrote:\n>> I timed a few common operations with both a warm OS file cache and a cold\n>> cache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did\n>> the operation in question a few times (first timing is the cold timing,\n>> the next few are the warm timings).  The following results are on a server\n>> with average hard drive (I.e., not flash)  and>  10GB of ram.\n>>\n>> 'git status' :   39 minutes cold, and 24 seconds warm.\n>>\n>> 'git blame':   44 minutes cold, 11 minutes warm.\n>>\n>> 'git add' (appending a few chars to the end of a file and adding it):   7\n>> seconds cold and 5 seconds warm.\n>>\n>> 'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n>> --no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\n>> of git to remove the three or four places where 'git commit' stats every\n>> file in the repo, and this dropped the times to 30 minutes cold and 8\n>> seconds warm.\n> Have you tried \"git update-index --assume-unchaged\"? That should\n> reduce mass lstat() and hopefully improve the above numbers. The\n> interface is not exactly easy-to-use, but if it has significant gain,\n> then we can try to improve UI.\n>\n> On the index size issue, ideally we should make minimum writes to\n> index instead of rewriting 191 MB index. An improvement we could do\n> now is to compress it, reduce disk footprint, thus disk I/O. If you\n> compress the index with gzip, how big is it?\nIf you're not afraid to add filesystem-specific code to git, you could \nleverage the btrfs find-new command (or use the ioctl directly) to \nquickly find changed files since a certain point in time. Other CoW \nfilesystems may have similar mechanisms. You could for example store the \nlast generation id in an index extension, that's what those extensions \nare for, right?\n\ntom\n"},{"id":"183900","messageId":"CACsJy8BRRWnPhO6U_WrfXa_pr0R0Zm9m=ZRgVjpUERYkELjxwA@mail.gmail.com","threadId":"29536","inReplyTo":"4F2E99C2.7090609@dbservice.com","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-05T15:17:00Z","receivedAt":"2012-02-05T15:17:00Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Feb 5, 2012 at 10:01 PM, Tomas Carnecky <tom@dbservice.com> wrote:\n> On 2/4/12 7:53 AM, Nguyen Thai Ngoc Duy wrote:\n>>\n>> On Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone<joshua.redstone@fb.com>\n>>  wrote:\n>>>\n>>> I timed a few common operations with both a warm OS file cache and a cold\n>>> cache.  i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then\n>>> did\n>>> the operation in question a few times (first timing is the cold timing,\n>>> the next few are the warm timings).  The following results are on a\n>>> server\n>>> with average hard drive (I.e., not flash)  and>  10GB of ram.\n>>>\n>>> 'git status' :   39 minutes cold, and 24 seconds warm.\n>>>\n>>> 'git blame':   44 minutes cold, 11 minutes warm.\n>>>\n>>> 'git add' (appending a few chars to the end of a file and adding it):   7\n>>> seconds cold and 5 seconds warm.\n>>>\n>>> 'git commit -m \"foo bar3\" --no-verify --untracked-files=no --quiet\n>>> --no-status':  41 minutes cold, 20 seconds warm.  I also hacked a version\n>>> of git to remove the three or four places where 'git commit' stats every\n>>> file in the repo, and this dropped the times to 30 minutes cold and 8\n>>> seconds warm.\n>>\n>> Have you tried \"git update-index --assume-unchaged\"? That should\n>> reduce mass lstat() and hopefully improve the above numbers. The\n>> interface is not exactly easy-to-use, but if it has significant gain,\n>> then we can try to improve UI.\n>>\n>> On the index size issue, ideally we should make minimum writes to\n>> index instead of rewriting 191 MB index. An improvement we could do\n>> now is to compress it, reduce disk footprint, thus disk I/O. If you\n>> compress the index with gzip, how big is it?\n>\n> If you're not afraid to add filesystem-specific code to git, you could\n> leverage the btrfs find-new command (or use the ioctl directly) to quickly\n> find changed files since a certain point in time. Other CoW filesystems may\n> have similar mechanisms. You could for example store the last generation id\n> in an index extension, that's what those extensions are for, right?\n\nSure they could be stored as index extensions. I'm more concerned of\nthe index size. I guess fs-specific code, if properly implemented\n(e.g. clean, handling repos crossing fs boundaries, moving repos...),\nmay get Junio's approval. There were also talks of implementing NTFS's\njournal (or something) on msysgit for similar goal.\n-- \nDuy\n"},{"id":"183991","messageId":"loom.20120206T054853-321@post.gmane.org","threadId":"29536","inReplyTo":"243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com","subject":"Re: Git performance results on a large repository","fromName":"David Mohs","fromEmail":"dgma@mohsinc.com","sentAt":"2012-02-06T07:10:13Z","receivedAt":"2012-02-06T07:10:13Z","isPatch":false,"sender":{"key":"dgma@mohsinc.com","avatar":null},"body":"Joshua Redstone <joshua.redstone <at> fb.com> writes:\n\n> To get a bit abstract for a moment, in an ideal world, it doesn't seem like\n> performance constraints of a source-control-system should dictate how we\n> choose to structure our code. Ideally, seems like we should be able to choose\n> to structure our code in whatever way we feel maximizes developer\n> productivity. If development and code/release management seem easier in a\n> single repo, than why not make an SCM that can handle it? This is one reason\n> I've been leaning towards figuring out an SCM approach that can work well with\n> our current practices rather than changing them as a prerequisite for good SCM\n> performance.\n\nI certainly agree with this perspective---that our tools should support our\nuse cases and not the other way around. However, I'd like you to consider that\nthe size of this hypothetical repository might be giving you some useful\ninformation on the health of the code it contains. You might consider creating\nseparate repositories simply to promote good modularization. It would involve\nsome up-front effort and certainly some pain, but this work itself might be\nbeneficial to your codebase without even considering the improved performance\nof the version control system.\n\nMy concern here is that it may be extremely difficult to make a single piece\nof software scale for a project that can grow arbitrarily large. You may add\nsome great performance improvements to git to then find that your bottleneck\nis the filesystem. That would enlarge the scope of your work and would likely\nmake the project more difficult to manage.\n\nIf you are able to prove me wrong, the entire software community will benefit\nfrom this work. However, before you embark upon a technical solution to your\nproblem, I would urge you to consider the possible benefits of a non-technical\nsolution, specifically restructuring your code and/or teams into more\nindependent modules. You might find benefits from this approach that extend\nbeyond source code control, which could make it the solution with the least\namount of overall risk.\n\nThanks for starting this valuable discussion.\n\n-David\n"},{"id":"184028","messageId":"20120206154043.GA14632@gnu.kitenet.net","threadId":"29536","inReplyTo":"CACsJy8Bf95JMp1qOiruR7+Tdi7JN42KNeMqGLud+z3O26DREnw@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Joey Hess","fromEmail":"joey@kitenet.net","sentAt":"2012-02-06T15:40:43Z","receivedAt":"2012-02-06T15:40:43Z","isPatch":false,"sender":{"key":"joey@kitenet.net","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Nguyen Thai Ngoc Duy wrote:\n> The \"interface to report which files have changed\" is exactly \"git\n> update-index --[no-]assume-unchanged\" is for. Have a look at the man\n> page. Basically you can mark every file \"unchanged\" in the beginning\n> and git won't bother lstat() them. What files you change, you have to\n> explicitly run \"git update-index --no-assume-unchanged\" to tell git.\n> \n> Someone on HN suggested making assume-unchanged files read-only to\n> avoid 90% accidentally changing a file without telling git. When\n> assume-unchanged bit is cleared, the file is made read-write again.\n\nThat made me think about using assume-unchanged with git-annex since it\nalready has read-only files. \n\nBut, here's what seems a misfeature... If an assume-unstaged file has\nmodifications and I git add it, nothing happens. To stage a change, I\nhave to explicitly git update-index --no-assume-unchanged and only then\ngit add, and then I need to remember to reset the assume-unstaged bit\nwhen I'm done working on that file for now. Compare with running git mv\non the same file, which does stage the move despite assume-unstaged. (So\ndoes git rm.)\n\n-- \nsee shy jo\n"},{"id":"184031","messageId":"CALts4TRGj1_uPX2b86GyfHHcDAUp6JSSMGmKjfS0p79DSAZ_uA@mail.gmail.com","threadId":"29536","inReplyTo":"243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com","subject":"Re: Git performance results on a large repository","fromName":"Matt Graham","fromEmail":"mdg149@gmail.com","sentAt":"2012-02-06T16:23:58Z","receivedAt":"2012-02-06T16:23:58Z","isPatch":false,"sender":{"key":"mdg149@gmail.com","avatar":"https://gravatar.com/avatar/a1f130a60a6550f75e8d7d3849e58e46494f36bfaf764a38cfd695ac85de8576?d=mp&s=160"},"body":"On Sat, Feb 4, 2012 at 18:05, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> [ wanted to reply to my initial msg, but wasn't subscribed to the list at time of mailing, so replying to most recent post instead ]\n>\n> Matt Graham:  I don't have file stats at the moment.  It's mostly code files, with a few larger data files here and there.    We also don't do sparse checkouts, primarily because most people use git (whether on top of SVN or not), which doesn't support it.\n\n\nThis doesn't help your original goal, but while you're still working\nwith git-svn, you can do sparse checkouts. Use --ignore-paths when you\ndo the original clone and it will filter out directories that are not\nof interest.\n\nWe used this at Etsy to keep git svn checkouts manageable when we\nstill had a gigantic svn repo.  You've repeatedly said you don't want\nto reorganize your repos but you may find this writeup informative\nabout how Etsy migrated to git (which included a health amount of repo\nmanipuation).\nhttp://codeascraft.etsy.com/2011/12/02/moving-from-svn-to-git-in-1000-easy-steps/\n"},{"id":"184055","messageId":"CB55A6A4.40AFD%joshua.redstone@fb.com","threadId":"29536","inReplyTo":"CALts4TRGj1_uPX2b86GyfHHcDAUp6JSSMGmKjfS0p79DSAZ_uA@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2012-02-06T20:50:08Z","receivedAt":"2012-02-06T20:50:08Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi all,\n\nNguyen, thanks for pointing out the assume-unchanged part.  That, and\nespecially the suggestion of making assume-unchanged files read-only is\ninteresting.  It does require explicit specification of what's changed.\nHmm, I wonder if that could be a candidate API through which something\nlike  CoW file system could let git know what's changed.  Btw, I think you\nasked earlier, but the index compresses from 158MB to 58MB - keep in mind\nthat the majority of file names in the repo are synthetic, so take with\nbig grain of salt.\n\nJoey, it sounds like it might be good if git-mv and other commands where\nconsistent in how they treat the assume-unchanged bit.\n\nDavid Mohs:  Yeah, it's an open question whether we'd be better off\nsomehow forcing the repos the split apart more.  As a practical matter,\nwhat may happen is that we incrementally solve our problem by addressing\npain points as they come up (e.g., git status being slow).  One risk with\nthat approach is that it leads to overly short-term thinking and we get\nstuck in a local minimum.  I totally agree that good modularization and\ncode health is valuable.  I think sometimes that getting to good\nmodularization does involve some technical work - like maybe moving\nfunctionality between systems so they split apart better, having some\nnotion of versioning and dependency and managing that, and so forth.    I\nsuppose the other aspect to the problem is that we want to make sure we\nhave a good source-control story even if the modularization effort takes a\nlong time - we'd rather not end up in a race between long-term\nmodularization efforts and source-control performance going south too\nfast.  I suppose this comes back to the desire that modularization not be\na prerequisite for good source-control performance.  Oh, and in case I\ndidn't mention it - we are working on modularization and splitting off\nlarge chunks of code, both into separable libraries as well as into\nseparate services, but it's a long-term process.\n\nMatt, some of our repos are still on SVN, many are on pure-git.  One of\nthe main ones that is on SVN is, at least at the moment, not amenable to\nsparse checkouts because of it's structure.\n\nTomas, yeah, I think one of the big questions is how little technical work\ncan we get away with, and where's the point of maximum leverage in terms\nof how much engineering time we invest.\n\nGreg,  'git commit' does some stat'ing of every file, even with all those\nflags - for example, I think one instance it does it is, just in case any\npre-commit hooks touched any files, it re-stats everything.  Regarding the\nperf numbers, I ran it on a beefy linux box.  Have you tried doing your\nmeasurements with the drop_caches trick to make sure the file cache is\ntotally cold?  Sorry for the dumb question, but how do I check the vnode\ncache size?\n\nDavid Lang and David Barr, I generated the pack files by doing a repack:\n\"git repack -a -d -f --max-pack-size=10g --depth=100 --window=250\"  after\ngenerating the repo.\n\nOne other update, the command I was running to get a histogram of all\nfiles in the repo finally completed.  The histogram (counting file size in\nbytes) is:\n\n[       0.0 -        6.4): 3\n[       6.4 -       41.3): 27\n[      41.3 -      265.7): 6\n[     265.7 -     1708.1): 652594\n[    1708.1 -    10980.6): 673482\n[   10980.6 -    70591.6): 19519\n[   70591.6 -   453814.3): 1583\n[  453814.3 -  2917451.4): 276\n[ 2917451.4 - 18755519.0): 61\n[18755519.0 - 120574242.0]: 4\nn=1347555 mean=3697.917708, median=1770.000000, stddev=122940.890559\n\nThe smaller files are all text (code), and the large ones are probably\nbinary.\n\nCheers,\nJosh\n\n\n\nOn 2/6/12 11:23 AM, \"Matt Graham\" <mdg149@gmail.com> wrote:\n\n>On Sat, Feb 4, 2012 at 18:05, Joshua Redstone <joshua.redstone@fb.com>\n>wrote:\n>> [ wanted to reply to my initial msg, but wasn't subscribed to the list\n>>at time of mailing, so replying to most recent post instead ]\n>>\n>> Matt Graham:  I don't have file stats at the moment.  It's mostly code\n>>files, with a few larger data files here and there.    We also don't do\n>>sparse checkouts, primarily because most people use git (whether on top\n>>of SVN or not), which doesn't support it.\n>\n>\n>This doesn't help your original goal, but while you're still working\n>with git-svn, you can do sparse checkouts. Use --ignore-paths when you\n>do the original clone and it will filter out directories that are not\n>of interest.\n>\n>We used this at Etsy to keep git svn checkouts manageable when we\n>still had a gigantic svn repo.  You've repeatedly said you don't want\n>to reorganize your repos but you may find this writeup informative\n>about how Etsy migrated to git (which included a health amount of repo\n>manipuation).\n>http://codeascraft.etsy.com/2011/12/02/moving-from-svn-to-git-in-1000-easy\n>-steps/\n"},{"id":"184058","messageId":"rmir4y7vff5.fsf@fnord.ir.bbn.com","threadId":"29536","inReplyTo":"CB55A6A4.40AFD%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Greg Troxel","fromEmail":"gdt@ir.bbn.com","sentAt":"2012-02-06T21:07:58Z","receivedAt":"2012-02-06T21:07:58Z","isPatch":false,"sender":{"key":"gdt@ir.bbn.com","avatar":null},"body":"\nJoshua Redstone <joshua.redstone@fb.com> writes:\n\n> Greg,  'git commit' does some stat'ing of every file, even with all those\n> flags - for example, I think one instance it does it is, just in case any\n> pre-commit hooks touched any files, it re-stats everything.\n\nThat seems ripe for skipping.  If I understand correctly, what's being\ncommitted is the index, not the working dir contents, so it would follow\nthat a pre-commit hook changing a file is a bug.\n\n> Regarding the perf numbers, I ran it on a beefy linux box.  Have you\n> tried doing your measurements with the drop_caches trick to make sure\n> the file cache is totally cold?\n\nOn NetBSD, there should be a clear cache command for just this reason,\nbut I'm not sure there is.  So I did\n\n  sysctl -w kern.maxvnodes=1000 # seemed to take a while\n  ls -lR # wait for those to be faulted in\n  sysctl -w kern.maxvnodes=500000\n\nThen, git status on my repo churned the disk for a long time.\n\n  real    2m7.121s\n  user    0m3.086s\n  sys     0m7.577s\n\nand then again right away\n\n  real    0m6.497s\n  user    0m2.533s\n  sys     0m3.010s\n\nThat repo has 217852 files (a real source tree with a few binaries, not\nsynthetic).\n\n> Sorry for the dumb question, but how do I check the vnode cache size?\n\nOn BSD, sysctl kern.maxvnodes.  I would aasume that on Linux there is\nsome max size for the the vnode cache, and that stat of a file in that\ncache is faster than going to the filesystem (even if reading from\ncached disk blocks).  But I really don't know how that works in Linux.\n\nI was going to say that if your vnode cache isn't big enough, then the\nhot run won't be so much faster than the warm run, but that's not true,\nbecause the fs blocks will be in the block cache and it will still help.\n"},{"id":"184061","messageId":"4F30435B.5070709@vilain.net","threadId":"29536","inReplyTo":"243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com","subject":"Re: Git performance results on a large repository","fromName":"Sam Vilain","fromEmail":"sam@vilain.net","sentAt":"2012-02-06T21:17:15Z","receivedAt":"2012-02-06T21:17:15Z","isPatch":false,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":" > Sam Vilain: Thanks for the pointer, i didn't realize that\n > fast-import was bi-directional.  I used it for generating the\n > synthetic repo.  Will look into using it the other way around.\n > Though that still won't speed up things like git-blame,\n > presumably?\n\nIt could, because blame is an operation which primarily works on\nthe source history with little reference to the working copy.  Of\ncourse this will depend on the quality of the implementation\nserver-side.  Blame should suit distribution over a cluster, as\nit is mostly involved with scanning candidate revisions for\nstring matches which is the compute intensive part.  Coming up\nwith candidate revisions has its own cost and can probably also\nbe distributed, but just working on the lowest loop level might\nbe a good place to start.\n\nWhat it doesn't help with is local filesystem operations.  For\nthis I think a different approach is required, if you can tie\ninto fam or a similar inode change notification system, then you\nshould be able to avoid the entire recursive stat on 'git\nstatus'.  I'm not sure --assume-unchanged on its own is a good\nidea, you could easily miss things.  Those stat's are useful.\n\nMaking the index able to hold just changes to the checked-out\ntree, as others have mentioned, would also save the massive reads\nand writes you've identified.  Perhaps a more high performance\nback-end could be developed.\n\n > The sparse-checkout issue you mention is a good one.\n\nIt's actually been on the table since at least GitTogether 2008;\nthere's been some design discussion on it and I think it's just\none of those features which doesn't have enough demand yet for it\nto be built.  It keeps coming up but not from anyone with the\ninclination or resources to make it happen.  There is a protocol\nissue, but this should be able to fit into the current extension\nsystem.\n\n > There is a good question of how to support quick checkout,\n > branch switching, clone, push and so forth.\n\nSure.  It will be much more network intensive as you are\nreplacing the part which normally has a very fast link through\nthe buffercache to pack files etc.  A hybrid approach is also\npossible, where objects are fetched individually via fast-import\nand cached in a local .git repo.  And I have a hunch that LZOP\ncompression of the stream may also be a win, but as with all of\nthese ideas, it would be after profiling identifies it as a choke point \nthan just because it sounds good.\n\n > I'll look into the approaches you suggest.  One consideration\n > is coming up with a high-leverage approach - i.e. not doing\n > heavy dev work if we can avoid it.\n\nRight.  You don't actually need to port the whole of git to Hadoop \ninitially, to begin with it can just pass through all commands to a \nserver-side git fast-import process.  When you find specific operations \nwhich are slow then these specific operations can be implemented using a \nHadoop back-end, and the rest backed to the standard git.  If done using \na useful plug-in system, these systems could be accepted by the core \nproject as an enterprise scaling option.\n\nThis could let you get going with the knowledge that the scaling option \nis there should it come out.\n\n > On the other hand, it would be nice if we (including the entire\n > community:) ) improve git in areas that others that share\n > similar issues benefit from as well.\n\nLike I say, a lot of people have run into this already...\n\nHTH,\nSam\n"},{"id":"184085","messageId":"CACsJy8D_yT3wzX1+Yfnwn7mtPiXz1smDGXxCtW62gHcCnTt0mw@mail.gmail.com","threadId":"29536","inReplyTo":"4F2C6276.1070100@vilain.net","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-07T01:19:47Z","receivedAt":"2012-02-07T01:19:47Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, Feb 4, 2012 at 5:40 AM, Sam Vilain <sam@vilain.net> wrote:\n> There have also been designs at various times for sparse check–outs; ie\n> check–outs where you don't check out the root of the repository but a\n> sub–tree.\n\nThere is a sparse checkout feature in git (hopefully from one of the\ndesigns you mentioned) and it can checkout subtrees. The only problem\nin this case is it maintains full index. So it only solves half of the\nproblem (stat calls), reading/writing large index just slows\neverything down.\n-- \nDuy\n"},{"id":"184086","messageId":"alpine.DEB.2.02.1202061722310.1107@asgard.lang.hm","threadId":"29536","inReplyTo":"CB55A6A4.40AFD%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"","fromEmail":"david@lang.hm","sentAt":"2012-02-07T01:28:06Z","receivedAt":"2012-02-07T01:28:06Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 6 Feb 2012, Joshua Redstone wrote:\n\n> David Lang and David Barr, I generated the pack files by doing a repack:\n> \"git repack -a -d -f --max-pack-size=10g --depth=100 --window=250\"  after\n> generating the repo.\n\nhow many pack files does this end up creating?\n\nI think that doing a full repack the way you did will group all revisions \nof a given file into a pack.\n\nwhile what I'm saying is that if you create the packs based on time, \nrather than space efficiency of the resulting pack files, you may end up \nnot having to go through as much date when doing things like a git blame.\n\nwhat you did was\n\ninitialize repo\n4M commits\nrepack\n\nwhat I'm saying is\n\ninitialize repo\nloop\n    500K commits\n    repack (and set pack to .keep so it doesn't get overwritten)\n\nso you will end up with ~8 sets of pack files, but time based so that when \nyou only need recent information you only look at the most recent pack \nfile. If you need to go back through all time, the multiple pack files \nwill be a little more expensive to process.\n\nthis has the added advantage that the 8 small repacks should be cheaper \nthan the one large repack as it isn't trying to cover all commits each \ntime.\n\nDavid Lang\n"},{"id":"184111","messageId":"loom.20120207T095317-899@post.gmane.org","threadId":"29536","inReplyTo":"CB5074CF.3AD7A%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Emanuele Zattin","fromEmail":"emanuelez@gmail.com","sentAt":"2012-02-07T08:58:02Z","receivedAt":"2012-02-07T08:58:02Z","isPatch":false,"sender":{"key":"emanuelez@gmail.com","avatar":null},"body":"Joshua Redstone <joshua.redstone <at> fb.com> writes:\n\n> \n> Hi Git folks,\n> \n\nHello everybody! \n\nI would just like to contribute a small set of blog posts \nabout this issue and a possible solution. \nSorry for the tone in which I wrote those posts, \nbut I think there are some valid points in there.\n\nhttps://gist.github.com/1758346\n\nBR,\n\nEmanuele Zattin\n"},{"id":"184130","messageId":"CACsJy8AxOZQ7S42V1g-b0vdBxPpjhFZe6qDkGaALnxQ6LiUssw@mail.gmail.com","threadId":"29536","inReplyTo":"20120206154043.GA14632@gnu.kitenet.net","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-07T13:43:50Z","receivedAt":"2012-02-07T13:43:50Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Mon, Feb 6, 2012 at 10:40 PM, Joey Hess <joey@kitenet.net> wrote:\n>> Someone on HN suggested making assume-unchanged files read-only to\n>> avoid 90% accidentally changing a file without telling git. When\n>> assume-unchanged bit is cleared, the file is made read-write again.\n>\n> That made me think about using assume-unchanged with git-annex since it\n> already has read-only files.\n>\n> But, here's what seems a misfeature...\n\nbecause, well.. assume-unchanged was designed to avoid stat() and\nnothing else. We are basing a new feature on top of it.\n\n> If an assume-unstaged file has\n> modifications and I git add it, nothing happens. To stage a change, I\n> have to explicitly git update-index --no-assume-unchanged and only then\n> git add, and then I need to remember to reset the assume-unstaged bit\n> when I'm done working on that file for now. Compare with running git mv\n> on the same file, which does stage the move despite assume-unstaged. (So\n> does git rm.)\n\nThis is normal in the lock-based \"checkout/edit/checkin\" model. mv/rm\noperates on directory content, which is not \"locked - no edit allowed\"\n(in our case --assume-unchanged) in git. But lock-based model does not\nmap really well to git anyway. It does not have the index (which may\nmake things more complicated). Also at index level, git does not\nreally understand directories.\n\nI think we could add a protection layer to index, where any changes\n(including removal) to an index entry are only allowed if the entry is\n\"unlocked\" (i.e no assume-unchanged bit). Locked entries are read-only\nand have assume-unchanged bit set. \"git (un)lock\" are introduced as\nnew UI. Does that make assume-unchanged friendlier?\n-- \nDuy\n"},{"id":"184266","messageId":"CB599BA0.42A6B%joshua.redstone@fb.com","threadId":"29536","inReplyTo":"CACsJy8AxOZQ7S42V1g-b0vdBxPpjhFZe6qDkGaALnxQ6LiUssw@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2012-02-09T21:06:21Z","receivedAt":"2012-02-09T21:06:21Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi Nguyen,\nI like the notion of using --assume-unchanged to cut down the set of\nthings that git considers may have changed.\nIt seems to me that there may still be situations that require operations\non the order of the # of files in the repo and hence may still be slow.\nFollowing is a list of potential candidates that occur to me.\n\n1. Switching branches, especially if you switch to an old branch.\nSometimes I've seen branch switching taking a long time for what I thought\nwas close to where HEAD was.\n\n2. Interactive rebase in which you reorder a few commits close to the tip\nof the branch (I observed this taking a long time, but haven't profiled it\nyet).  I include here other types of cherry-picking of commits.\n\n3. Any working directory operations that fail part-way through and make\nyou want to do a 'git reset --hard' or at least a full 'git-status'.  That\nis, when you have reason to believe that files with 'assume-unchange' may\nhave accidentally changed.\n\n4. Operations that require rewriting the index - I think git-add is one?\n\nIf the working-tree representation is the full set of all files\nmaterialized on disk and it's the same as the representation of files\nchanged, then I'm not sure how to avoid some of these without playing file\nsystem games or using wrapper scripts.\n\nWhat do you (or others) think?\n\n\nJosh\n\n\nOn 2/7/12 8:43 AM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n\n>On Mon, Feb 6, 2012 at 10:40 PM, Joey Hess <joey@kitenet.net> wrote:\n>>> Someone on HN suggested making assume-unchanged files read-only to\n>>> avoid 90% accidentally changing a file without telling git. When\n>>> assume-unchanged bit is cleared, the file is made read-write again.\n>>\n>> That made me think about using assume-unchanged with git-annex since it\n>> already has read-only files.\n>>\n>> But, here's what seems a misfeature...\n>\n>because, well.. assume-unchanged was designed to avoid stat() and\n>nothing else. We are basing a new feature on top of it.\n>\n>> If an assume-unstaged file has\n>> modifications and I git add it, nothing happens. To stage a change, I\n>> have to explicitly git update-index --no-assume-unchanged and only then\n>> git add, and then I need to remember to reset the assume-unstaged bit\n>> when I'm done working on that file for now. Compare with running git mv\n>> on the same file, which does stage the move despite assume-unstaged. (So\n>> does git rm.)\n>\n>This is normal in the lock-based \"checkout/edit/checkin\" model. mv/rm\n>operates on directory content, which is not \"locked - no edit allowed\"\n>(in our case --assume-unchanged) in git. But lock-based model does not\n>map really well to git anyway. It does not have the index (which may\n>make things more complicated). Also at index level, git does not\n>really understand directories.\n>\n>I think we could add a protection layer to index, where any changes\n>(including removal) to an index entry are only allowed if the entry is\n>\"unlocked\" (i.e no assume-unchanged bit). Locked entries are read-only\n>and have assume-unchanged bit set. \"git (un)lock\" are introduced as\n>new UI. Does that make assume-unchanged friendlier?\n>-- \n>Duy\n>--\n>To unsubscribe from this list: send the line \"unsubscribe git\" in\n>the body of a message to majordomo@vger.kernel.org\n>More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"184329","messageId":"CACsJy8DQNHm8sTgxKL=+Ui5OBsJBpvPn+dRmN9bVMwq4TfNuxQ@mail.gmail.com","threadId":"29536","inReplyTo":"CB599BA0.42A6B%joshua.redstone@fb.com","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-10T07:12:47Z","receivedAt":"2012-02-10T07:12:47Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Feb 10, 2012 at 4:06 AM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> Hi Nguyen,\n> I like the notion of using --assume-unchanged to cut down the set of\n> things that git considers may have changed.\n> It seems to me that there may still be situations that require operations\n> on the order of the # of files in the repo and hence may still be slow.\n> Following is a list of potential candidates that occur to me.\n>\n> 1. Switching branches, especially if you switch to an old branch.\n> Sometimes I've seen branch switching taking a long time for what I thought\n> was close to where HEAD was.\n>\n> 2. Interactive rebase in which you reorder a few commits close to the tip\n> of the branch (I observed this taking a long time, but haven't profiled it\n> yet).  I include here other types of cherry-picking of commits.\n>\n> 3. Any working directory operations that fail part-way through and make\n> you want to do a 'git reset --hard' or at least a full 'git-status'.  That\n> is, when you have reason to believe that files with 'assume-unchange' may\n> have accidentally changed.\n\nAll these involve unpack_trees(), which is full tree operation. The\nbigger your worktree is, the slower it is. Another good reason to\nsplit unrelated parts into separate repositories.\n\n\n> 4. Operations that require rewriting the index - I think git-add is one?\n>\n> If the working-tree representation is the full set of all files\n> materialized on disk and it's the same as the representation of files\n> changed, then I'm not sure how to avoid some of these without playing file\n> system games or using wrapper scripts.\n>\n> What do you (or others) think?\n>\n>\n> Josh\n>\n>\n> On 2/7/12 8:43 AM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n>\n>>On Mon, Feb 6, 2012 at 10:40 PM, Joey Hess <joey@kitenet.net> wrote:\n>>>> Someone on HN suggested making assume-unchanged files read-only to\n>>>> avoid 90% accidentally changing a file without telling git. When\n>>>> assume-unchanged bit is cleared, the file is made read-write again.\n>>>\n>>> That made me think about using assume-unchanged with git-annex since it\n>>> already has read-only files.\n>>>\n>>> But, here's what seems a misfeature...\n>>\n>>because, well.. assume-unchanged was designed to avoid stat() and\n>>nothing else. We are basing a new feature on top of it.\n>>\n>>> If an assume-unstaged file has\n>>> modifications and I git add it, nothing happens. To stage a change, I\n>>> have to explicitly git update-index --no-assume-unchanged and only then\n>>> git add, and then I need to remember to reset the assume-unstaged bit\n>>> when I'm done working on that file for now. Compare with running git mv\n>>> on the same file, which does stage the move despite assume-unstaged. (So\n>>> does git rm.)\n>>\n>>This is normal in the lock-based \"checkout/edit/checkin\" model. mv/rm\n>>operates on directory content, which is not \"locked - no edit allowed\"\n>>(in our case --assume-unchanged) in git. But lock-based model does not\n>>map really well to git anyway. It does not have the index (which may\n>>make things more complicated). Also at index level, git does not\n>>really understand directories.\n>>\n>>I think we could add a protection layer to index, where any changes\n>>(including removal) to an index entry are only allowed if the entry is\n>>\"unlocked\" (i.e no assume-unchanged bit). Locked entries are read-only\n>>and have assume-unchanged bit set. \"git (un)lock\" are introduced as\n>>new UI. Does that make assume-unchanged friendlier?\n>>--\n>>Duy\n>>--\n>>To unsubscribe from this list: send the line \"unsubscribe git\" in\n>>the body of a message to majordomo@vger.kernel.org\n>>More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \nDuy\n"},{"id":"184340","messageId":"CAP8UFD1RTa6+btjJrsfqjOOoCjebZBqK6xkPN7ZVLM04bHO9yw@mail.gmail.com","threadId":"29536","inReplyTo":"CACsJy8DQNHm8sTgxKL=+Ui5OBsJBpvPn+dRmN9bVMwq4TfNuxQ@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2012-02-10T09:39:56Z","receivedAt":"2012-02-10T09:39:56Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi,\n\nOn Fri, Feb 10, 2012 at 8:12 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n>\n> All these involve unpack_trees(), which is full tree operation. The\n> bigger your worktree is, the slower it is. Another good reason to\n> split unrelated parts into separate repositories.\n\nMaybe having different \"views\" would be enough to make a smaller\nworktree and history, so that things are much faster for a developper?\n\n(I already suggested \"views\" based on \"git replace\" in this thread:\nhttp://thread.gmane.org/gmane.comp.version-control.git/177146/focus=177639)\n\nBest regards,\nChristian.\n"},{"id":"184346","messageId":"CACsJy8ANHdG5r10Hk4Ap74+=KGrtDJGYbQa+9731S2YeHCb2Yw@mail.gmail.com","threadId":"29536","inReplyTo":"CAP8UFD1RTa6+btjJrsfqjOOoCjebZBqK6xkPN7ZVLM04bHO9yw@mail.gmail.com","subject":"Re: Git performance results on a large repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-10T12:24:55Z","receivedAt":"2012-02-10T12:24:55Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Feb 10, 2012 at 4:39 PM, Christian Couder\n<christian.couder@gmail.com> wrote:\n> Hi,\n>\n> On Fri, Feb 10, 2012 at 8:12 AM, Nguyen Thai Ngoc Duy <pclouds@gmail.com> wrote:\n>>\n>> All these involve unpack_trees(), which is full tree operation. The\n>> bigger your worktree is, the slower it is. Another good reason to\n>> split unrelated parts into separate repositories.\n>\n> Maybe having different \"views\" would be enough to make a smaller\n> worktree and history, so that things are much faster for a developper?\n>\n> (I already suggested \"views\" based on \"git replace\" in this thread:\n> http://thread.gmane.org/gmane.comp.version-control.git/177146/focus=177639)\n\nThat's more or less what I did with the subtree clone series [1] and\nended up doing narrow clone [2]. The only difference between the two\nare how to handle partial worktree/index. The former uses git-replace\nto seal any holes, the latter tackles at pathspec level and is\ngenerally more elegant.\n\nThe worktree part from that work should be usable in full clone too. I\nam reviving the series and going to repost it soon. Have a look [3] if\nyou are interested.\n\n[1] http://thread.gmane.org/gmane.comp.version-control.git/152347\n[2] http://thread.gmane.org/gmane.comp.version-control.git/155427\n[3] https://github.com/pclouds/git/commits/narrow-clone\n-- \nDuy\n"}]}