{"thread":{"id":"29073","subject":"Debugging git-commit slowness on a large repo","startedAt":"2011-12-02T23:17:10Z","lastAt":"2011-12-20T19:26:50Z","messageCount":16,"participants":["Joshua Redstone","Carlos Martín Nieto","Tomas Carnecky","Junio C Hamano","Nguyen Thai Ngoc Duy","Thomas Rast"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"180272","messageId":"CAFE9C7B.2BFEC%joshua.redstone@fb.com","threadId":"29073","inReplyTo":null,"subject":"Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-02T23:17:10Z","receivedAt":"2011-12-02T23:17:10Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi,\nI have a git repo with about 300k commits,  150k files totaling maybe 7GB.\n Locally committing a small change - say touching fewer than 300 bytes\nacross 4 files - consistently takes over one second, which seems kinda\nslow.  This is using git 1.7.7.4 on a linux 2.6 box.  The time does not\nimprove after doing a git-gc (my .git dir has maybe 250 files after a git\ngc).  The same size commit on a brand new repo takes < 10ms.  Any thoughts\non why committing a small change seems to take a long time on larger repos?\n\nFwiw, I also tried doing the same test using libgit2 (via the pygit2\nwrapper), and it was ever slower (about 6 seconds to commit the same small\nchange).\n\nThanks for any thoughts or places to look.\n\nCheers,\nJosh\n"},{"id":"180275","messageId":"20111203002347.GB2950@centaur.lab.cmartin.tk","threadId":"29073","inReplyTo":"CAFE9C7B.2BFEC%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Carlos Martín Nieto","fromEmail":"cmn@elego.de","sentAt":"2011-12-03T00:23:47Z","receivedAt":"2011-12-03T00:23:47Z","isPatch":false,"sender":{"key":"cmn@elego.de","avatar":"https://avatars.githubusercontent.com/u/335443?v=4"},"body":"On Fri, Dec 02, 2011 at 11:17:10PM +0000, Joshua Redstone wrote:\n> Hi,\n> I have a git repo with about 300k commits,  150k files totaling maybe 7GB.\n>  Locally committing a small change - say touching fewer than 300 bytes\n> across 4 files - consistently takes over one second, which seems kinda\n> slow.  This is using git 1.7.7.4 on a linux 2.6 box.  The time does not\n> improve after doing a git-gc (my .git dir has maybe 250 files after a git\n> gc).  The same size commit on a brand new repo takes < 10ms.  Any thoughts\n> on why committing a small change seems to take a long time on larger repos?\n\nBy \"same size commit\" do you mean the same amount of changes, or the\nsame amount of files? Committing doesn't depend on the size of the\nrepo (by itself), but on the size of the index, which depends on the\namount of files to be committed (as git is snapshot-based). At one\npoint, commit forgot how to write the tree cache to the index (a\nperformance optimisation). Do the times improve if you run 'git\nread-tree HEAD' between one commit and another? Note that this will\nreset the index to the last commit, though for the tests I image you\nuse some variation of 'git commit -a'.\n\nThomas Rast wrote a patch to re-teach commit to store the tree cache,\nbut there were some issues and never got applied.\n\n> \n> Fwiw, I also tried doing the same test using libgit2 (via the pygit2\n> wrapper), and it was ever slower (about 6 seconds to commit the same small\n> change).\n\nI don't know about the python bindings, but on the (somewhat\nunscientific) tests for libgit2's write-tree (the slow part of a\ncreating a commit), it performs slightly faster than git's (though I\nthink git's write-tree does update the tree cache, which libgit2\ndoesn't currently). The speed could just be a side-effect of the small\ntest repo. From your domain, I assume the data is not for public\nconsumption, but it'd be great if you could post your code to pygit2's\nissue tracker so we can see how much of the slowdown comes from the\nbindings or the library.\n\n   cmn\n\n"},{"id":"180293","messageId":"4EDB7B82.8090707@dbservice.com","threadId":"29073","inReplyTo":"CAFE9C7B.2BFEC%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Tomas Carnecky","fromEmail":"tom@dbservice.com","sentAt":"2011-12-04T13:54:10Z","receivedAt":"2011-12-04T13:54:10Z","isPatch":false,"sender":{"key":"tom@dbservice.com","avatar":"https://gravatar.com/avatar/900a300bdd1a8bbe086008ad78210bbee2ad2803b7d50a5cba04c1e9404bd6d2?d=mp&s=160"},"body":"On 12/3/11 12:17 AM, Joshua Redstone wrote:\n> Hi,\n> I have a git repo with about 300k commits,  150k files totaling maybe 7GB.\n>   Locally committing a small change - say touching fewer than 300 bytes\n> across 4 files - consistently takes over one second, which seems kinda\n> slow.  This is using git 1.7.7.4 on a linux 2.6 box.  The time does not\n> improve after doing a git-gc (my .git dir has maybe 250 files after a git\n> gc).  The same size commit on a brand new repo takes<  10ms.  Any thoughts\n> on why committing a small change seems to take a long time on larger repos?\n>\n> Fwiw, I also tried doing the same test using libgit2 (via the pygit2\n> wrapper), and it was ever slower (about 6 seconds to commit the same small\n> change).\n\ntry git commit --no-status\n"},{"id":"180307","messageId":"7vr50inckx.fsf@alter.siamese.dyndns.org","threadId":"29073","inReplyTo":"20111203002347.GB2950@centaur.lab.cmartin.tk","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-12-05T17:38:22Z","receivedAt":"2011-12-05T17:38:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Carlos Martín Nieto <cmn@elego.de> writes:\n\n> ... At one\n> point, commit forgot how to write the tree cache to the index (a\n> performance optimisation). Do the times improve if you run 'git\n> read-tree HEAD' between one commit and another? Note that this will\n> reset the index to the last commit, though for the tests I image you\n> use some variation of 'git commit -a'.\n>\n> Thomas Rast wrote a patch to re-teach commit to store the tree cache,\n> but there were some issues and never got applied.\n\nAhh, I forgot all about that exchange.\n\n  http://thread.gmane.org/gmane.comp.version-control.git/178480/focus=178515\n\nThe cache-tree mechanism has traditionally been one of the more important\noptimizations and it would be very nice if we can resurrect the behaviour\nfor \"git commit\" too.\n"},{"id":"180437","messageId":"CB04005C.2C669%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"20111203002347.GB2950@centaur.lab.cmartin.tk","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-07T01:48:46Z","receivedAt":"2011-12-07T01:48:46Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi Carlos and Tomas and Junio,\n\n@Tomas, I tried adding the '--no-status' flag to 'git commit' and it sped\nthings up by maybe 15%, but commits still take a second.\n\n@Carlos, by \"same size\", I mean roughly the same number of files and\nnumber of bytes modified in each file.  In all experiments, it's less than\n5 files modified per commit with changes totaling fewer than 10 KB, often\nmore like 1 KB.  I actually wrote a test script to generate commits,\ncustomized for the stats on the repo I'm using.  It repeatedly generates\nsome changes, does 'git add [ list of files changed ]' followed by 'git\ncommit --no-status -m [ msg ]'.   It generates changes by picking fewer\nthan 5 files at random, modifying two 100-byte regions in each file, and\noccasionally creates a new file of about 1 KB.  If it helps, I can\nprobably post the test script I've been using.\n\nI tried doing a 'git read-tree HEAD' before each 'git add ; git commit'\niteration, and the time for git-commit jumped from about 1 second to about\n8 seconds.  That is a pretty dramatic slowdown.  Any idea why?  I wonder\nif that's related to the overall commit slowness.\n\n@Carlos and/or @Junio, can you point me at any docs/code to understand\nwhat a tree-cache is and how it differs from the index?  I did a google\nsearch for [git tree-cache index], but nothing popped out.\n\nCheers,\nJosh\n\n\nOn 12/2/11 4:23 PM, \"Carlos Martín Nieto\" <cmn@elego.de> wrote:\n\n>On Fri, Dec 02, 2011 at 11:17:10PM +0000, Joshua Redstone wrote:\n>> Hi,\n>> I have a git repo with about 300k commits,  150k files totaling maybe\n>>7GB.\n>>  Locally committing a small change - say touching fewer than 300 bytes\n>> across 4 files - consistently takes over one second, which seems kinda\n>> slow.  This is using git 1.7.7.4 on a linux 2.6 box.  The time does not\n>> improve after doing a git-gc (my .git dir has maybe 250 files after a\n>>git\n>> gc).  The same size commit on a brand new repo takes < 10ms.  Any\n>>thoughts\n>> on why committing a small change seems to take a long time on larger\n>>repos?\n>\n>By \"same size commit\" do you mean the same amount of changes, or the\n>same amount of files? Committing doesn't depend on the size of the\n>repo (by itself), but on the size of the index, which depends on the\n>amount of files to be committed (as git is snapshot-based). At one\n>point, commit forgot how to write the tree cache to the index (a\n>performance optimisation). Do the times improve if you run 'git\n>read-tree HEAD' between one commit and another? Note that this will\n>reset the index to the last commit, though for the tests I image you\n>use some variation of 'git commit -a'.\n>\n>Thomas Rast wrote a patch to re-teach commit to store the tree cache,\n>but there were some issues and never got applied.\n>\n>> \n>> Fwiw, I also tried doing the same test using libgit2 (via the pygit2\n>> wrapper), and it was ever slower (about 6 seconds to commit the same\n>>small\n>> change).\n>\n>I don't know about the python bindings, but on the (somewhat\n>unscientific) tests for libgit2's write-tree (the slow part of a\n>creating a commit), it performs slightly faster than git's (though I\n>think git's write-tree does update the tree cache, which libgit2\n>doesn't currently). The speed could just be a side-effect of the small\n>test repo. From your domain, I assume the data is not for public\n>consumption, but it'd be great if you could post your code to pygit2's\n>issue tracker so we can see how much of the slowdown comes from the\n>bindings or the library.\n>\n>   cmn\n>\n"},{"id":"180438","messageId":"CACsJy8Dbd+v+8FzvQS9a4C8DQSxQGgqQNGaLhL1cHv-yMnaCJQ@mail.gmail.com","threadId":"29073","inReplyTo":"CB04005C.2C669%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-12-07T02:08:55Z","receivedAt":"2011-12-07T02:08:55Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Dec 7, 2011 at 8:48 AM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> I tried doing a 'git read-tree HEAD' before each 'git add ; git commit'\n> iteration, and the time for git-commit jumped from about 1 second to about\n> 8 seconds.  That is a pretty dramatic slowdown.  Any idea why?  I wonder\n> if that's related to the overall commit slowness.\n\nHow big is your working directory? \"git ls-files | wc -l\" should show\nit. Try \"git read-tree HEAD; git add; git write-tree\" and see if the\nwrite-tree part takes as much time as commit. write-tree is mainly\nabout cache-tree generation.\n\n> @Carlos and/or @Junio, can you point me at any docs/code to understand\n> what a tree-cache is and how it differs from the index?  I did a google\n> search for [git tree-cache index], but nothing popped out.\n\nHave a look at Documentation/technical/index-format.txt. Cache tree\nextension is near the end.\n-- \nDuy\n"},{"id":"180539","messageId":"CB051EFC.2C795%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"CACsJy8Dbd+v+8FzvQS9a4C8DQSxQGgqQNGaLhL1cHv-yMnaCJQ@mail.gmail.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-07T22:48:22Z","receivedAt":"2011-12-07T22:48:22Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Hi Duy,\nThanks for the documentation link.\n\ngit ls-files shows 100k files, which matches # of files in the working\ntree ('find . -type f -print | wc -l').\n\nI added a 'git read-tree HEAD' before the git-add, and a 'git write-tree'\nafter the add.  With that, the commit time slowed down to 8 seconds per\ncommit, plus 4 more seconds for the read-tree/add/write-tree ops.  The\nread-tree/add/write-tree each took about a second.\n\nAs an experiment, I also tried removing the 'git read-tree' and just\nhaving the git-write-tree.  That sped up commits to 0.6 seconds, but the\noverall time for add/write-tree/commit was still 3 to 6 seconds.\n\nFor comparison, without the read-tree and write-tree, commits take about 1\nsecond and add/commit in total takes about 2 seconds.\n\nIt surprises me that the presence of git read-tree or write-tree would\nslow things down so much.\n\nJosh\n\nOn 12/6/11 6:08 PM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n\n>On Wed, Dec 7, 2011 at 8:48 AM, Joshua Redstone <joshua.redstone@fb.com>\n>wrote:\n>> I tried doing a 'git read-tree HEAD' before each 'git add ; git commit'\n>> iteration, and the time for git-commit jumped from about 1 second to\n>>about\n>> 8 seconds.  That is a pretty dramatic slowdown.  Any idea why?  I wonder\n>> if that's related to the overall commit slowness.\n>\n>How big is your working directory? \"git ls-files | wc -l\" should show\n>it. Try \"git read-tree HEAD; git add; git write-tree\" and see if the\n>write-tree part takes as much time as commit. write-tree is mainly\n>about cache-tree generation.\n>\n>> @Carlos and/or @Junio, can you point me at any docs/code to understand\n>> what a tree-cache is and how it differs from the index?  I did a google\n>> search for [git tree-cache index], but nothing popped out.\n>\n>Have a look at Documentation/technical/index-format.txt. Cache tree\n>extension is near the end.\n>-- \n>Duy\n"},{"id":"180549","messageId":"CACsJy8DiWWr7eo86gzb-XcqfDv4_ENkqWxswTNb-k84xO18c=A@mail.gmail.com","threadId":"29073","inReplyTo":"CB051EFC.2C795%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2011-12-08T01:39:32Z","receivedAt":"2011-12-08T01:39:32Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Dec 8, 2011 at 5:48 AM, Joshua Redstone <joshua.redstone@fb.com> wrote:\n> Hi Duy,\n> Thanks for the documentation link.\n>\n> git ls-files shows 100k files, which matches # of files in the working\n> tree ('find . -type f -print | wc -l').\n\nAny chance you can split it into smaller repositories, or remove files\nfrom working directory (e.g. if you store logs, you don't have to keep\nlogs from all time in working directory, they can be retrieved from\nhistory).\n\n> I added a 'git read-tree HEAD' before the git-add, and a 'git write-tree'\n> after the add.  With that, the commit time slowed down to 8 seconds per\n> commit, plus 4 more seconds for the read-tree/add/write-tree ops.  The\n> read-tree/add/write-tree each took about a second.\n\nread-tree destroys stat info in index, refreshing 100k entries in\nindex in this case may take some time. Try this to see if commit time\nreduces and how much time update-index takes\n\nread-tree HEAD\nupdate-index --refresh\nadd ....\nwrite-tree\ncommit -q\n\n> As an experiment, I also tried removing the 'git read-tree' and just\n> having the git-write-tree.  That sped up commits to 0.6 seconds, but the\n> overall time for add/write-tree/commit was still 3 to 6 seconds.\n\noverall time is not really important because we duplicate work here\n(write-tree is done as part of commit again). What I'm trying to do is\nto determine how much time each operation in commit may take.\n-- \nDuy\n"},{"id":"180624","messageId":"CB069000.2C9C6%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"CACsJy8DiWWr7eo86gzb-XcqfDv4_ENkqWxswTNb-k84xO18c=A@mail.gmail.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-09T00:09:52Z","receivedAt":"2011-12-09T00:09:52Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"On 12/7/11 5:39 PM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n\n>On Thu, Dec 8, 2011 at 5:48 AM, Joshua Redstone <joshua.redstone@fb.com>\n>wrote:\n>> Hi Duy,\n>> Thanks for the documentation link.\n>>\n>> git ls-files shows 100k files, which matches # of files in the working\n>> tree ('find . -type f -print | wc -l').\n>\n>Any chance you can split it into smaller repositories, or remove files\n>from working directory (e.g. if you store logs, you don't have to keep\n>logs from all time in working directory, they can be retrieved from\n>history).\n\nIt's not really feasible to split it into smaller repositories.  In fact,\nwe're expecting it to grow between 3x and 5x in number of files and number\nof commits.\n\n>\n>> I added a 'git read-tree HEAD' before the git-add, and a 'git\n>>write-tree'\n>> after the add.  With that, the commit time slowed down to 8 seconds per\n>> commit, plus 4 more seconds for the read-tree/add/write-tree ops.  The\n>> read-tree/add/write-tree each took about a second.\n>\n>read-tree destroys stat info in index, refreshing 100k entries in\n>index in this case may take some time. Try this to see if commit time\n>reduces and how much time update-index takes\n>\n>read-tree HEAD\n>update-index --refresh\n>add ....\n>write-tree\n>commit -q\n\nI added the \"update-index --refresh\" and the time for commit became more\nlike 0.6 seconds.\nIn this setup: read-tree takes ~2 seconds, update-index takes ~8 seconds,\ngit-add takes 1 to 4 seconds, and write-tree takes less than 1 second.\n\n>\n>> As an experiment, I also tried removing the 'git read-tree' and just\n>> having the git-write-tree.  That sped up commits to 0.6 seconds, but the\n>> overall time for add/write-tree/commit was still 3 to 6 seconds.\n>\n>overall time is not really important because we duplicate work here\n>(write-tree is done as part of commit again). What I'm trying to do is\n>to determine how much time each operation in commit may take.\n>-- \n>Duy\n"},{"id":"180625","messageId":"CB069308.2C9DD%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"CB069000.2C9C6%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-09T00:17:55Z","receivedAt":"2011-12-09T00:17:55Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Btw, I also tried doing some very poor-man's profiling on \"git commit\"\nwithout any of the readtree/writetree/updateindex commands.\n\nAround 50% of the time was in (bottom few frames may have varied)\n\n#1  0x00000000004c467e in find_pack_entry (sha1=0x1475a44 ,\ne=0x7fff2621f070) at sha1_file.c:2027\n#2  0x00000000004c57b0 in has_sha1_file (sha1=0x7fe2cd9c7900 \"00\") at\nsha1_file.c:2567   \n                   \n                 \n#3  0x000000000046e4af in update_one (it=<value optimized out>,\ncache=<value optimized out>, entries=<value optimized out>, base=<value\noptimized out>, baselen=<value optimized out>, missing_ok=<value optimized\nout>, dryrun=0) at cache-\\\ntree.c:333         \n                   \n                   \n            \n#4  0x000000000046e278 in update_one (it=<value optimized out>,\ncache=<value optimized out>, entries=<value optimized out>, base=<value\noptimized out>, baselen=<value optimized out>, missing_ok=<value optimized\nout>, dryrun=0) at cache-\\\ntree.c:285         \n                   \n                   \n            \n#5  0x000000000046e278 in update_one (it=<value optimized out>,\ncache=<value optimized out>, entries=<value optimized out>, base=<value\noptimized out>, baselen=<value optimized out>, missing_ok=<value optimized\nout>, dryrun=0) at cache-\\\ntree.c:285         \n                   \n                   \n            \n#6  0x000000000046e278 in update_one (it=<value optimized out>,\ncache=<value optimized out>, entries=<value optimized out>, base=<value\noptimized out>, baselen=<value optimized out>, missing_ok=<value optimized\nout>, dryrun=0) at cache-\\\ntree.c:285         \n                   \n                   \n            \n#7  0x000000000046e278 in update_one (it=<value optimized out>,\ncache=<value optimized out>, entries=<value optimized out>, base=<value\noptimized out>, baselen=<value optimized out>, missing_ok=<value optimized\nout>, dryrun=0) at cache-\\\ntree.c:285         \n                   \n                   \n            \n#8  0x000000000046e278 in update_one (it=<value optimized out>,\ncache=<value optimized out>, entries=<value optimized out>, base=<value\noptimized out>, baselen=<value optimized out>, missing_ok=<value optimized\nout>, dryrun=0) at cache-\\\ntree.c:285         \n                   \n                   \n            \n#9  0x000000000046e869 in cache_tree_update (it=<value optimized out>,\ncache=<value optimized out>, entries=dwarf2_read_address: Corrupted DWARF\nexpression.        \n                 \n) at cache-tree.c:379\n                   \n                   \n            \n#10 0x000000000041cade in prepare_to_commit (index_file=0x781740\n\".git/index\", prefix=<value optimized out>, current_head=<value optimized\nout>, s=0x7fff26220d00, author_ident=<value optimized out>) at\nbuiltin/commit.c:866\n#11 0x000000000041d891 in cmd_commit (argc=0, argv=0x7fff262213a0,\nprefix=0x0) at builtin/commit.c:1407\n                   \n                   \n#12 0x0000000000404bf7 in handle_internal_command (argc=4,\nargv=0x7fff262213a0) at git.c:308\n                   \n                   \n#13 0x0000000000404e2f in main (argc=4, argv=0x7fff262213a0) at git.c:512\n                   \n                   \n            \n \n\n\nAnd 30% of the time was in:\n\n#0  0x00000034af2c34a5 in _lxstat () from /lib64/libc.so.6\n                   \n                   \n            \n#1  0x00000000004abe0f in refresh_cache_ent (istate=0x780940,\nce=0x7f8462a34e40, options=0, err=0x7fff6dd9f588) at\n/usr/include/sys/stat.h:443\n                   \n#2  0x00000000004ac1a0 in refresh_index (istate=0x780940, flags=<value\noptimized out>, pathspec=<value optimized out>, seen=<value optimized\nout>, header_msg=0x0) at read-cache.c:1133\n                   \n#3  0x000000000041b60a in refresh_cache_or_die (refresh_flags=<value\noptimized out>) at builtin/commit.c:331\n                   \n                  \n#4  0x000000000041bc39 in prepare_index (argc=0, argv=0x7fff6dda0310,\nprefix=0x0, current_head=<value optimized out>, is_status=<value optimized\nout>) at builtin/commit.c:414\n                 \n#5  0x000000000041d878 in cmd_commit (argc=0, argv=0x7fff6dda0310,\nprefix=0x0) at builtin/commit.c:1403\n                   \n                   \n  \n\n\nJosh\n\n\nOn 12/8/11 4:09 PM, \"Joshua Redstone\" <joshua.redstone@fb.com> wrote:\n\n>On 12/7/11 5:39 PM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n>\n>>On Thu, Dec 8, 2011 at 5:48 AM, Joshua Redstone <joshua.redstone@fb.com>\n>>wrote:\n>>> Hi Duy,\n>>> Thanks for the documentation link.\n>>>\n>>> git ls-files shows 100k files, which matches # of files in the working\n>>> tree ('find . -type f -print | wc -l').\n>>\n>>Any chance you can split it into smaller repositories, or remove files\n>>from working directory (e.g. if you store logs, you don't have to keep\n>>logs from all time in working directory, they can be retrieved from\n>>history).\n>\n>It's not really feasible to split it into smaller repositories.  In fact,\n>we're expecting it to grow between 3x and 5x in number of files and number\n>of commits.\n>\n>>\n>>> I added a 'git read-tree HEAD' before the git-add, and a 'git\n>>>write-tree'\n>>> after the add.  With that, the commit time slowed down to 8 seconds per\n>>> commit, plus 4 more seconds for the read-tree/add/write-tree ops.  The\n>>> read-tree/add/write-tree each took about a second.\n>>\n>>read-tree destroys stat info in index, refreshing 100k entries in\n>>index in this case may take some time. Try this to see if commit time\n>>reduces and how much time update-index takes\n>>\n>>read-tree HEAD\n>>update-index --refresh\n>>add ....\n>>write-tree\n>>commit -q\n>\n>I added the \"update-index --refresh\" and the time for commit became more\n>like 0.6 seconds.\n>In this setup: read-tree takes ~2 seconds, update-index takes ~8 seconds,\n>git-add takes 1 to 4 seconds, and write-tree takes less than 1 second.\n>\n>>\n>>> As an experiment, I also tried removing the 'git read-tree' and just\n>>> having the git-write-tree.  That sped up commits to 0.6 seconds, but\n>>>the\n>>> overall time for add/write-tree/commit was still 3 to 6 seconds.\n>>\n>>overall time is not really important because we duplicate work here\n>>(write-tree is done as part of commit again). What I'm trying to do is\n>>to determine how much time each operation in commit may take.\n>>-- \n>>Duy\n>\n"},{"id":"180998","messageId":"CB0BCE02.2CD42%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"CB069308.2C9DD%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-13T00:15:30Z","receivedAt":"2011-12-13T00:15:30Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"Sorry for the poor formatting of the stack trace.\n\nI've written two scripts to reproduce the slow commit behavior that I see.\n I've posted both to:\n   https://gist.github.com/1469760\n\nTo repro, first create a dir with lots of files (it defaults to creating 1\nmillion files in 1000 dirs):\n\n$ loadGen.py --baseDir=./bigdir\n\nthen, run the simulator scripts to generate and commit a series of small\nchanges to the repo:\n\n$ git reset --hard HEAD && simulate.py ./bigdir git\n\nThe git reset is to clean up any cruft left over from a previous partial\ninvocation of simulate.py\n\nNote that loadGen.py defaults to creating 1 million files and committing\nthem in one commit.  With a flash drive this took < 30 min, and subsequent\nsmall commits in simulate.py took about 6 seconds.  With a hard-drive,\nit's taking > 1hr (still waiting for it to finish).\n\nCheers,\nJosh\n\n\nOn 12/8/11 4:17 PM, \"Joshua Redstone\" <joshua.redstone@fb.com> wrote:\n\n>Btw, I also tried doing some very poor-man's profiling on \"git commit\"\n>without any of the readtree/writetree/updateindex commands.\n>\n>Around 50% of the time was in (bottom few frames may have varied)\n>\n>#1  0x00000000004c467e in find_pack_entry (sha1=0x1475a44 ,\n>e=0x7fff2621f070) at sha1_file.c:2027\n>#2  0x00000000004c57b0 in has_sha1_file (sha1=0x7fe2cd9c7900 \"00\") at\n>sha1_file.c:2567  \n>                  \n>                 \n>#3  0x000000000046e4af in update_one (it=<value optimized out>,\n>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>optimized out>, baselen=<value optimized out>, missing_ok=<value optimized\n>out>, dryrun=0) at cache-\\\n>tree.c:333        \n>                  \n>                  \n>            \n>#4  0x000000000046e278 in update_one (it=<value optimized out>,\n>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>optimized out>, baselen=<value optimized out>, missing_ok=<value optimized\n>out>, dryrun=0) at cache-\\\n>tree.c:285        \n>                  \n>                  \n>            \n>#5  0x000000000046e278 in update_one (it=<value optimized out>,\n>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>optimized out>, baselen=<value optimized out>, missing_ok=<value optimized\n>out>, dryrun=0) at cache-\\\n>tree.c:285        \n>                  \n>                  \n>            \n>#6  0x000000000046e278 in update_one (it=<value optimized out>,\n>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>optimized out>, baselen=<value optimized out>, missing_ok=<value optimized\n>out>, dryrun=0) at cache-\\\n>tree.c:285        \n>                  \n>                  \n>            \n>#7  0x000000000046e278 in update_one (it=<value optimized out>,\n>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>optimized out>, baselen=<value optimized out>, missing_ok=<value optimized\n>out>, dryrun=0) at cache-\\\n>tree.c:285        \n>                  \n>                  \n>            \n>#8  0x000000000046e278 in update_one (it=<value optimized out>,\n>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>optimized out>, baselen=<value optimized out>, missing_ok=<value optimized\n>out>, dryrun=0) at cache-\\\n>tree.c:285        \n>                  \n>                  \n>            \n>#9  0x000000000046e869 in cache_tree_update (it=<value optimized out>,\n>cache=<value optimized out>, entries=dwarf2_read_address: Corrupted DWARF\n>expression.       \n>                 \n>) at cache-tree.c:379\n>                  \n>                  \n>            \n>#10 0x000000000041cade in prepare_to_commit (index_file=0x781740\n>\".git/index\", prefix=<value optimized out>, current_head=<value optimized\n>out>, s=0x7fff26220d00, author_ident=<value optimized out>) at\n>builtin/commit.c:866\n>#11 0x000000000041d891 in cmd_commit (argc=0, argv=0x7fff262213a0,\n>prefix=0x0) at builtin/commit.c:1407\n>                  \n>                  \n>#12 0x0000000000404bf7 in handle_internal_command (argc=4,\n>argv=0x7fff262213a0) at git.c:308\n>                  \n>                  \n>#13 0x0000000000404e2f in main (argc=4, argv=0x7fff262213a0) at git.c:512\n>                  \n>                  \n>            \n> \n>\n>\n>And 30% of the time was in:\n>\n>#0  0x00000034af2c34a5 in _lxstat () from /lib64/libc.so.6\n>                  \n>                  \n>            \n>#1  0x00000000004abe0f in refresh_cache_ent (istate=0x780940,\n>ce=0x7f8462a34e40, options=0, err=0x7fff6dd9f588) at\n>/usr/include/sys/stat.h:443\n>                  \n>#2  0x00000000004ac1a0 in refresh_index (istate=0x780940, flags=<value\n>optimized out>, pathspec=<value optimized out>, seen=<value optimized\n>out>, header_msg=0x0) at read-cache.c:1133\n>                  \n>#3  0x000000000041b60a in refresh_cache_or_die (refresh_flags=<value\n>optimized out>) at builtin/commit.c:331\n>                  \n>                  \n>#4  0x000000000041bc39 in prepare_index (argc=0, argv=0x7fff6dda0310,\n>prefix=0x0, current_head=<value optimized out>, is_status=<value optimized\n>out>) at builtin/commit.c:414\n>                 \n>#5  0x000000000041d878 in cmd_commit (argc=0, argv=0x7fff6dda0310,\n>prefix=0x0) at builtin/commit.c:1403\n>                  \n>                  \n>  \n>\n>\n>Josh\n>\n>\n>On 12/8/11 4:09 PM, \"Joshua Redstone\" <joshua.redstone@fb.com> wrote:\n>\n>>On 12/7/11 5:39 PM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n>>\n>>>On Thu, Dec 8, 2011 at 5:48 AM, Joshua Redstone <joshua.redstone@fb.com>\n>>>wrote:\n>>>> Hi Duy,\n>>>> Thanks for the documentation link.\n>>>>\n>>>> git ls-files shows 100k files, which matches # of files in the working\n>>>> tree ('find . -type f -print | wc -l').\n>>>\n>>>Any chance you can split it into smaller repositories, or remove files\n>>>from working directory (e.g. if you store logs, you don't have to keep\n>>>logs from all time in working directory, they can be retrieved from\n>>>history).\n>>\n>>It's not really feasible to split it into smaller repositories.  In fact,\n>>we're expecting it to grow between 3x and 5x in number of files and\n>>number\n>>of commits.\n>>\n>>>\n>>>> I added a 'git read-tree HEAD' before the git-add, and a 'git\n>>>>write-tree'\n>>>> after the add.  With that, the commit time slowed down to 8 seconds\n>>>>per\n>>>> commit, plus 4 more seconds for the read-tree/add/write-tree ops.  The\n>>>> read-tree/add/write-tree each took about a second.\n>>>\n>>>read-tree destroys stat info in index, refreshing 100k entries in\n>>>index in this case may take some time. Try this to see if commit time\n>>>reduces and how much time update-index takes\n>>>\n>>>read-tree HEAD\n>>>update-index --refresh\n>>>add ....\n>>>write-tree\n>>>commit -q\n>>\n>>I added the \"update-index --refresh\" and the time for commit became more\n>>like 0.6 seconds.\n>>In this setup: read-tree takes ~2 seconds, update-index takes ~8 seconds,\n>>git-add takes 1 to 4 seconds, and write-tree takes less than 1 second.\n>>\n>>>\n>>>> As an experiment, I also tried removing the 'git read-tree' and just\n>>>> having the git-write-tree.  That sped up commits to 0.6 seconds, but\n>>>>the\n>>>> overall time for add/write-tree/commit was still 3 to 6 seconds.\n>>>\n>>>overall time is not really important because we duplicate work here\n>>>(write-tree is done as part of commit again). What I'm trying to do is\n>>>to determine how much time each operation in commit may take.\n>>>-- \n>>>Duy\n>>\n>\n"},{"id":"181502","messageId":"CB1518AB.2D649%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"CB0BCE02.2CD42%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-20T00:51:16Z","receivedAt":"2011-12-20T00:51:16Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"I've managed to speed up git-commit on large repos by 4x by removing some\nsafeguards that caused git to stat every file in the repo on commits that\ntouch a small number of files.  The diff, for illustrative purposes only,\nis at:\n\n    https://gist.github.com/1499621\n\n\nWith a repo with 1 million files (but few commits), the diff drops the\ncommit time down from 7.3 seconds to 1.8 seconds, a 75% decrease. The\noptimizations are:\n\n1. Remove call to refresh_cache_or_die that stats every file in the repo,\ni think the purpose is to detect any changes between git-add and\ngit-commit.\n\n2. Pass missing_ok=true to cache_tree_update. This causes the tree\ngeneration code to not stat every file in the repo to verify it still\nexists as a git object.\n\n3. Remove pair discard_cache/read_cache_from, which rereads the index\nfile. I think this was in case a pre-commit hook changed the set of things\nbeing committed.\n\nIt may be worth making some of these flag-enabled.\n\n\n\nJosh\n\n\nOn 12/12/11 4:15 PM, \"Joshua Redstone\" <joshua.redstone@fb.com> wrote:\n\n>Sorry for the poor formatting of the stack trace.\n>\n>I've written two scripts to reproduce the slow commit behavior that I see.\n> I've posted both to:\n>   https://gist.github.com/1469760\n>\n>To repro, first create a dir with lots of files (it defaults to creating 1\n>million files in 1000 dirs):\n>\n>$ loadGen.py --baseDir=./bigdir\n>\n>then, run the simulator scripts to generate and commit a series of small\n>changes to the repo:\n>\n>$ git reset --hard HEAD && simulate.py ./bigdir git\n>\n>The git reset is to clean up any cruft left over from a previous partial\n>invocation of simulate.py\n>\n>Note that loadGen.py defaults to creating 1 million files and committing\n>them in one commit.  With a flash drive this took < 30 min, and subsequent\n>small commits in simulate.py took about 6 seconds.  With a hard-drive,\n>it's taking > 1hr (still waiting for it to finish).\n>\n>Cheers,\n>Josh\n>\n>\n>On 12/8/11 4:17 PM, \"Joshua Redstone\" <joshua.redstone@fb.com> wrote:\n>\n>>Btw, I also tried doing some very poor-man's profiling on \"git commit\"\n>>without any of the readtree/writetree/updateindex commands.\n>>\n>>Around 50% of the time was in (bottom few frames may have varied)\n>>\n>>#1  0x00000000004c467e in find_pack_entry (sha1=0x1475a44 ,\n>>e=0x7fff2621f070) at sha1_file.c:2027\n>>#2  0x00000000004c57b0 in has_sha1_file (sha1=0x7fe2cd9c7900 \"00\") at\n>>sha1_file.c:2567 \n>>                 \n>>                 \n>>#3  0x000000000046e4af in update_one (it=<value optimized out>,\n>>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>>optimized out>, baselen=<value optimized out>, missing_ok=<value\n>>optimized\n>>out>, dryrun=0) at cache-\\\n>>tree.c:333       \n>>                 \n>>                 \n>>            \n>>#4  0x000000000046e278 in update_one (it=<value optimized out>,\n>>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>>optimized out>, baselen=<value optimized out>, missing_ok=<value\n>>optimized\n>>out>, dryrun=0) at cache-\\\n>>tree.c:285       \n>>                 \n>>                 \n>>            \n>>#5  0x000000000046e278 in update_one (it=<value optimized out>,\n>>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>>optimized out>, baselen=<value optimized out>, missing_ok=<value\n>>optimized\n>>out>, dryrun=0) at cache-\\\n>>tree.c:285       \n>>                 \n>>                 \n>>            \n>>#6  0x000000000046e278 in update_one (it=<value optimized out>,\n>>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>>optimized out>, baselen=<value optimized out>, missing_ok=<value\n>>optimized\n>>out>, dryrun=0) at cache-\\\n>>tree.c:285       \n>>                 \n>>                 \n>>            \n>>#7  0x000000000046e278 in update_one (it=<value optimized out>,\n>>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>>optimized out>, baselen=<value optimized out>, missing_ok=<value\n>>optimized\n>>out>, dryrun=0) at cache-\\\n>>tree.c:285       \n>>                 \n>>                 \n>>            \n>>#8  0x000000000046e278 in update_one (it=<value optimized out>,\n>>cache=<value optimized out>, entries=<value optimized out>, base=<value\n>>optimized out>, baselen=<value optimized out>, missing_ok=<value\n>>optimized\n>>out>, dryrun=0) at cache-\\\n>>tree.c:285       \n>>                 \n>>                 \n>>            \n>>#9  0x000000000046e869 in cache_tree_update (it=<value optimized out>,\n>>cache=<value optimized out>, entries=dwarf2_read_address: Corrupted DWARF\n>>expression.      \n>>                 \n>>) at cache-tree.c:379\n>>                 \n>>                 \n>>            \n>>#10 0x000000000041cade in prepare_to_commit (index_file=0x781740\n>>\".git/index\", prefix=<value optimized out>, current_head=<value optimized\n>>out>, s=0x7fff26220d00, author_ident=<value optimized out>) at\n>>builtin/commit.c:866\n>>#11 0x000000000041d891 in cmd_commit (argc=0, argv=0x7fff262213a0,\n>>prefix=0x0) at builtin/commit.c:1407\n>>                 \n>>                 \n>>#12 0x0000000000404bf7 in handle_internal_command (argc=4,\n>>argv=0x7fff262213a0) at git.c:308\n>>                 \n>>                 \n>>#13 0x0000000000404e2f in main (argc=4, argv=0x7fff262213a0) at git.c:512\n>>                 \n>>                 \n>>            \n>> \n>>\n>>\n>>And 30% of the time was in:\n>>\n>>#0  0x00000034af2c34a5 in _lxstat () from /lib64/libc.so.6\n>>                 \n>>                 \n>>            \n>>#1  0x00000000004abe0f in refresh_cache_ent (istate=0x780940,\n>>ce=0x7f8462a34e40, options=0, err=0x7fff6dd9f588) at\n>>/usr/include/sys/stat.h:443\n>>                 \n>>#2  0x00000000004ac1a0 in refresh_index (istate=0x780940, flags=<value\n>>optimized out>, pathspec=<value optimized out>, seen=<value optimized\n>>out>, header_msg=0x0) at read-cache.c:1133\n>>                 \n>>#3  0x000000000041b60a in refresh_cache_or_die (refresh_flags=<value\n>>optimized out>) at builtin/commit.c:331\n>>                 \n>>                 \n>>#4  0x000000000041bc39 in prepare_index (argc=0, argv=0x7fff6dda0310,\n>>prefix=0x0, current_head=<value optimized out>, is_status=<value\n>>optimized\n>>out>) at builtin/commit.c:414\n>>                 \n>>#5  0x000000000041d878 in cmd_commit (argc=0, argv=0x7fff6dda0310,\n>>prefix=0x0) at builtin/commit.c:1403\n>>                 \n>>                 \n>>  \n>>\n>>\n>>Josh\n>>\n>>\n>>On 12/8/11 4:09 PM, \"Joshua Redstone\" <joshua.redstone@fb.com> wrote:\n>>\n>>>On 12/7/11 5:39 PM, \"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> wrote:\n>>>\n>>>>On Thu, Dec 8, 2011 at 5:48 AM, Joshua Redstone\n>>>><joshua.redstone@fb.com>\n>>>>wrote:\n>>>>> Hi Duy,\n>>>>> Thanks for the documentation link.\n>>>>>\n>>>>> git ls-files shows 100k files, which matches # of files in the\n>>>>>working\n>>>>> tree ('find . -type f -print | wc -l').\n>>>>\n>>>>Any chance you can split it into smaller repositories, or remove files\n>>>>from working directory (e.g. if you store logs, you don't have to keep\n>>>>logs from all time in working directory, they can be retrieved from\n>>>>history).\n>>>\n>>>It's not really feasible to split it into smaller repositories.  In\n>>>fact,\n>>>we're expecting it to grow between 3x and 5x in number of files and\n>>>number\n>>>of commits.\n>>>\n>>>>\n>>>>> I added a 'git read-tree HEAD' before the git-add, and a 'git\n>>>>>write-tree'\n>>>>> after the add.  With that, the commit time slowed down to 8 seconds\n>>>>>per\n>>>>> commit, plus 4 more seconds for the read-tree/add/write-tree ops.\n>>>>>The\n>>>>> read-tree/add/write-tree each took about a second.\n>>>>\n>>>>read-tree destroys stat info in index, refreshing 100k entries in\n>>>>index in this case may take some time. Try this to see if commit time\n>>>>reduces and how much time update-index takes\n>>>>\n>>>>read-tree HEAD\n>>>>update-index --refresh\n>>>>add ....\n>>>>write-tree\n>>>>commit -q\n>>>\n>>>I added the \"update-index --refresh\" and the time for commit became more\n>>>like 0.6 seconds.\n>>>In this setup: read-tree takes ~2 seconds, update-index takes ~8\n>>>seconds,\n>>>git-add takes 1 to 4 seconds, and write-tree takes less than 1 second.\n>>>\n>>>>\n>>>>> As an experiment, I also tried removing the 'git read-tree' and just\n>>>>> having the git-write-tree.  That sped up commits to 0.6 seconds, but\n>>>>>the\n>>>>> overall time for add/write-tree/commit was still 3 to 6 seconds.\n>>>>\n>>>>overall time is not really important because we duplicate work here\n>>>>(write-tree is done as part of commit again). What I'm trying to do is\n>>>>to determine how much time each operation in commit may take.\n>>>>-- \n>>>>Duy\n>>>\n>>\n>\n"},{"id":"181504","messageId":"7vehw0kphc.fsf@alter.siamese.dyndns.org","threadId":"29073","inReplyTo":"CB1518AB.2D649%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-12-20T01:21:03Z","receivedAt":"2011-12-20T01:21:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Joshua Redstone <joshua.redstone@fb.com> writes:\n\n> I've managed to speed up git-commit on large repos by 4x by removing some\n> safeguards that caused git to stat every file in the repo on commits that\n> touch a small number of files.  The diff, for illustrative purposes only,\n> is at:\n>\n>     https://gist.github.com/1499621\n>\n>\n> With a repo with 1 million files (but few commits), the diff drops the\n> commit time down from 7.3 seconds to 1.8 seconds, a 75% decrease. The\n> optimizations are:\n\nI do not know if these kind of changes are called \"optimizations\" or\nmerely making the command record a random tree object that may have some\nresemblance to what you wanted to commit but is subtly incorrect. I didn't\nfetch your safety removal, though.\n\nWouldn't you get a similar speed-up without being unsafe if you simply ran\n\"git commit\" without any parameter (i.e. write out the current index as a\ntree and make a commit), combined with \"--no-status\" and perhaps \"-q\" to\navoid running the comparison between the resulting commit and the working\ntree state after the commit?\n"},{"id":"181509","messageId":"CB152498.2D6DB%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"7vehw0kphc.fsf@alter.siamese.dyndns.org","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-20T01:40:47Z","receivedAt":"2011-12-20T01:40:47Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"You're right, more than optimizations, they are modifications that reduce\nsafety checks and make assumptions about the way one is using git (e.g.,\nyou always remember to add each file you want to commit).  I focused on\nthem because:\n\n  1. In our installation, we don't use commit hooks that change what's\nbeing committed, so it's good to know that in principle, there's a big\nperf benefit to be had by leveraging that fact.\n\n  2. At an abstract level, it seems like the cost of doing a commit should\nbe proportional to the amount of the repository touched by the commit, not\nby the size of the repository.  These experiments are demonstrations of\none direction that a set of optimizations would need to go to get commit\nperformance more along those lines.\n\n  3. We're also exploring storage systems that support more efficient ways\nto query what's changed than stat'ing every file.\n\nI forgot to mention, the times I quoted where with --no-verify and\n--no-status.  Adding '-q' didn't speed up performance at all.\n\n\nAs a bonus, I've also profiled git-add on the 1-million file repo, and it\nlooks like, as you might expect, the time is dominated by reading and\nwriting the index.  The time for git-add is a couple of seconds.\n\nJosh\n\n\nOn 12/19/11 5:21 PM, \"Junio C Hamano\" <gitster@pobox.com> wrote:\n\n>Joshua Redstone <joshua.redstone@fb.com> writes:\n>\n>> I've managed to speed up git-commit on large repos by 4x by removing\n>>some\n>> safeguards that caused git to stat every file in the repo on commits\n>>that\n>> touch a small number of files.  The diff, for illustrative purposes\n>>only,\n>> is at:\n>>\n>>     https://gist.github.com/1499621\n>>\n>>\n>> With a repo with 1 million files (but few commits), the diff drops the\n>> commit time down from 7.3 seconds to 1.8 seconds, a 75% decrease. The\n>> optimizations are:\n>\n>I do not know if these kind of changes are called \"optimizations\" or\n>merely making the command record a random tree object that may have some\n>resemblance to what you wanted to commit but is subtly incorrect. I didn't\n>fetch your safety removal, though.\n>\n>Wouldn't you get a similar speed-up without being unsafe if you simply ran\n>\"git commit\" without any parameter (i.e. write out the current index as a\n>tree and make a commit), combined with \"--no-status\" and perhaps \"-q\" to\n>avoid running the comparison between the resulting commit and the working\n>tree state after the commit?\n"},{"id":"181520","messageId":"87wr9rk35n.fsf@thomas.inf.ethz.ch","threadId":"29073","inReplyTo":"CB152498.2D6DB%joshua.redstone@fb.com","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2011-12-20T09:23:16Z","receivedAt":"2011-12-20T09:23:16Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"Joshua Redstone <joshua.redstone@fb.com> writes:\n> As a bonus, I've also profiled git-add on the 1-million file repo, and it\n> looks like, as you might expect, the time is dominated by reading and\n> writing the index.  The time for git-add is a couple of seconds.\n\nNote that the time to write the index itself is also rather small, but\nthe time needed to sha1 the index when loading and then again when\nsaving it really hurts.\n\n(I noticed this while working on the commit-tree topic.)\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"181530","messageId":"CB162065.2E009%joshua.redstone@fb.com","threadId":"29073","inReplyTo":"87wr9rk35n.fsf@thomas.inf.ethz.ch","subject":"Re: Debugging git-commit slowness on a large repo","fromName":"Joshua Redstone","fromEmail":"joshua.redstone@fb.com","sentAt":"2011-12-20T19:26:50Z","receivedAt":"2011-12-20T19:26:50Z","isPatch":false,"sender":{"key":"joshua.redstone@fb.com","avatar":null},"body":"I looked again at my poor-mans-profiling output of git-add.  The Sha1\nstuff under ce_write_entry->ce_write_flush  takes a bunch of time.\ncommit_lock_file->rename takes about the same as well.\n\nBtw, the perf numbers for commit and add are with a warm file cache.  I\nexpect the benefit of skipping all the stat() calls will increase for cold\ncache.\n\nJosh\n\nOn 12/20/11 1:23 AM, \"Thomas Rast\" <trast@student.ethz.ch> wrote:\n\n>Joshua Redstone <joshua.redstone@fb.com> writes:\n>> As a bonus, I've also profiled git-add on the 1-million file repo, and\n>>it\n>> looks like, as you might expect, the time is dominated by reading and\n>> writing the index.  The time for git-add is a couple of seconds.\n>\n>Note that the time to write the index itself is also rather small, but\n>the time needed to sha1 the index when loading and then again when\n>saving it really hurts.\n>\n>(I noticed this while working on the commit-tree topic.)\n>\n>-- \n>Thomas Rast\n>trast@{inf,student}.ethz.ch\n"}]}