{"thread":{"id":"11182","subject":"Something is broken in repack","startedAt":"2007-12-07T23:05:38Z","lastAt":"2007-12-14T16:45:07Z","messageCount":82,"participants":["Jon Smirl","Linus Torvalds","Nicolas Pitre","David Brown","Harvey Harrison","Junio C Hamano","Morten Welinder","Sean","Andreas Ericsson","David Kastrup","Pierre Habouzit","David Miller","Daniel Berlin","Paolo Bonzini","Nguyen Thai Ngoc Duy","Johannes Sixt","Jakub Narebski","Wolfram Gloger"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"62346","messageId":"9e4733910712071505y6834f040k37261d65a2d445c4@mail.gmail.com","threadId":"11182","inReplyTo":null,"subject":"Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-07T23:05:38Z","receivedAt":"2007-12-07T23:05:38Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Using this config:\n[pack]\n        threads = 4\n        deltacachesize = 256M\n        deltacachelimit = 0\n\nAnd the 330MB gcc pack for input\n git repack -a -d -f  --depth=250 --window=250\n\ncomplete seconds RAM\n10%  47 1GB\n20%  29 1Gb\n30%  24 1Gb\n40%  18 1GB\n50%  110 1.2GB\n60%  85 1.4GB\n70%  195 1.5GB\n80%  186 2.5GB\n90%  489 3.8GB\n95%  800 4.8GB\nI killed it because it started swapping\n\nThe mmaps are only about 400MB in this case.\nAt the end the git process had 4.4GB of physical RAM allocated.\n\nStarting from a highly compressed pack greatly aggravates the problem.\nStarting with a 2GB pack of the same data my process size only grew to\n3GB with 2GB of mmaps.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62349","messageId":"alpine.LFD.0.9999.0712071632490.12046@woody.linux-foundation.org","threadId":"11182","inReplyTo":"9e4733910712071505y6834f040k37261d65a2d445c4@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-08T00:37:16Z","receivedAt":"2007-12-08T00:37:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 7 Dec 2007, Jon Smirl wrote:\n>\n> Using this config:\n> [pack]\n>         threads = 4\n>         deltacachesize = 256M\n\nI think deltacachesize is broken.\n\nThe code in try_delta() that replaces a delta cache entry with another one \nseems very buggy wrt that whole \"delta_cache_size\" update. It does\n\n\tdelta_cache_size -= trg_entry->delta_size;\n\nto account for the old delta going away, but it does this *after* having \nalready replaced trg_entry->delta_size with the new delta entry.\n\nI suspect there are other issues going on too, but that's the one that I \nnoticed from a quick look-through.\n\nNico? I think this one is yours..\n\n\t\tLinus\n"},{"id":"62352","messageId":"alpine.LFD.0.99999.0712072016590.555@xanadu.home","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712071632490.12046@woody.linux-foundation.org","subject":"[PATCH] pack-objects: fix delta cache size accounting","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-08T01:27:52Z","receivedAt":"2007-12-08T01:27:52Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"The wrong value was substracted from delta_cache_size when replacing\na cached delta, as trg_entry->delta_size was used after the old size\nhad been replaced by the new size.\n\nNoticed by Linus.\n\nSigned-off-by: Nicolas Pitre <nico@cam.org> \n---\n\nOn Fri, 7 Dec 2007, Linus Torvalds wrote:\n\n> The code in try_delta() that replaces a delta cache entry with another one \n> seems very buggy wrt that whole \"delta_cache_size\" update. It does\n> \n> \tdelta_cache_size -= trg_entry->delta_size;\n> \n> to account for the old delta going away, but it does this *after* having \n> already replaced trg_entry->delta_size with the new delta entry.\n\nDoh!  Mea culpa.\n\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 4f44658..350ece4 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -1422,10 +1422,6 @@ static int try_delta(struct unpacked *trg, struct unpacked *src,\n \t\t}\n \t}\n \n-\ttrg_entry->delta = src_entry;\n-\ttrg_entry->delta_size = delta_size;\n-\ttrg->depth = src->depth + 1;\n-\n \t/*\n \t * Handle memory allocation outside of the cache\n \t * accounting lock.  Compiler will optimize the strangeness\n@@ -1439,7 +1435,7 @@ static int try_delta(struct unpacked *trg, struct unpacked *src,\n \t\ttrg_entry->delta_data = NULL;\n \t}\n \tif (delta_cacheable(src_size, trg_size, delta_size)) {\n-\t\tdelta_cache_size += trg_entry->delta_size;\n+\t\tdelta_cache_size += delta_size;\n \t\tcache_unlock();\n \t\ttrg_entry->delta_data = xrealloc(delta_buf, delta_size);\n \t} else {\n@@ -1447,6 +1443,10 @@ static int try_delta(struct unpacked *trg, struct unpacked *src,\n \t\tfree(delta_buf);\n \t}\n \n+\ttrg_entry->delta = src_entry;\n+\ttrg_entry->delta_size = delta_size;\n+\ttrg->depth = src->depth + 1;\n+\n \treturn 1;\n }\n \n"},{"id":"62355","messageId":"alpine.LFD.0.99999.0712072032410.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712071505y6834f040k37261d65a2d445c4@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-08T01:46:25Z","receivedAt":"2007-12-08T01:46:25Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Dec 2007, Jon Smirl wrote:\n\n> Using this config:\n> [pack]\n>         threads = 4\n>         deltacachesize = 256M\n>         deltacachelimit = 0\n\nSince you have a different result according to the source pack used then \nthose cache settings, even if there was a bug with them, are not \nsignificant.\n\n> And the 330MB gcc pack for input\n>  git repack -a -d -f  --depth=250 --window=250\n> \n> complete seconds RAM\n> 10%  47 1GB\n> 20%  29 1Gb\n> 30%  24 1Gb\n> 40%  18 1GB\n> 50%  110 1.2GB\n> 60%  85 1.4GB\n> 70%  195 1.5GB\n> 80%  186 2.5GB\n> 90%  489 3.8GB\n> 95%  800 4.8GB\n> I killed it because it started swapping\n> \n> The mmaps are only about 400MB in this case.\n> At the end the git process had 4.4GB of physical RAM allocated.\n\nThat's really bad.\n\n> Starting from a highly compressed pack greatly aggravates the problem.\n\nThat is really interesting though.\n\n> Starting with a 2GB pack of the same data my process size only grew to\n> 3GB with 2GB of mmaps.\n\nWhich is quite reasonable, even if the same issue might still be there.\n\nSo the problem seems to be related to the pack access code and not the \nrepack code.  And it must have something to do with the number of deltas \nbeing replayed.  And because the repack is attempting delta compression \nroughly from newest to oldest, and because old objects are typically in \na deeper delta chain, then this might explain the logarithmic slowdown.\n\nSo something must be wrong with the delta cache in sha1_file.c somehow.\n\n\nNicolas\n"},{"id":"62357","messageId":"9e4733910712071804ja0a49e1m1eb209cb942bc36f@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712072032410.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-08T02:04:27Z","receivedAt":"2007-12-08T02:04:27Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/7/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 7 Dec 2007, Jon Smirl wrote:\n>\n> > Using this config:\n> > [pack]\n> >         threads = 4\n> >         deltacachesize = 256M\n> >         deltacachelimit = 0\n>\n> Since you have a different result according to the source pack used then\n> those cache settings, even if there was a bug with them, are not\n> significant.\n>\n> > And the 330MB gcc pack for input\n> >  git repack -a -d -f  --depth=250 --window=250\n> >\n> > complete seconds RAM\n> > 10%  47 1GB\n> > 20%  29 1Gb\n> > 30%  24 1Gb\n> > 40%  18 1GB\n> > 50%  110 1.2GB\n> > 60%  85 1.4GB\n> > 70%  195 1.5GB\n> > 80%  186 2.5GB\n> > 90%  489 3.8GB\n> > 95%  800 4.8GB\n> > I killed it because it started swapping\n> >\n> > The mmaps are only about 400MB in this case.\n> > At the end the git process had 4.4GB of physical RAM allocated.\n>\n> That's really bad.\n>\n> > Starting from a highly compressed pack greatly aggravates the problem.\n>\n> That is really interesting though.\n>\n> > Starting with a 2GB pack of the same data my process size only grew to\n> > 3GB with 2GB of mmaps.\n>\n> Which is quite reasonable, even if the same issue might still be there.\n>\n> So the problem seems to be related to the pack access code and not the\n> repack code.  And it must have something to do with the number of deltas\n> being replayed.  And because the repack is attempting delta compression\n> roughly from newest to oldest, and because old objects are typically in\n> a deeper delta chain, then this might explain the logarithmic slowdown.\n>\n> So something must be wrong with the delta cache in sha1_file.c somehow.\n\nI applied the delta accounting patch. It took about 200MB of from the\nmemory use but that doesn't make a dent in 4GB of allocations.\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62359","messageId":"9e4733910712071822w5a7d5bb5k5d099825b333acda@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712072032410.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-08T02:22:21Z","receivedAt":"2007-12-08T02:22:21Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/7/07, Nicolas Pitre <nico@cam.org> wrote:\n> So the problem seems to be related to the pack access code and not the\n> repack code.  And it must have something to do with the number of deltas\n> being replayed.  And because the repack is attempting delta compression\n> roughly from newest to oldest, and because old objects are typically in\n> a deeper delta chain, then this might explain the logarithmic slowdown.\n\nWhat could be wrongly allocating 4GB of memory? Figure that out and\nyou should have your answer. The slow down may be coming from having\nto search through more and more objects in memory.\n\nMemory consumption seem to be correlated to the depth of the delta\nchain being accessed. It blows up tremendously right at the end. It\nmay even be a square of the length of the chain length. For the normal\ndefault case the square didn't hurt, but 250*250 = 62,500 which would\neat a huge amount of memory.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62360","messageId":"alpine.LFD.0.99999.0712072124160.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712071804ja0a49e1m1eb209cb942bc36f@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-08T02:28:51Z","receivedAt":"2007-12-08T02:28:51Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Dec 2007, Jon Smirl wrote:\n\n> On 12/7/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Fri, 7 Dec 2007, Jon Smirl wrote:\n> >\n> > >  git repack -a -d -f  --depth=250 --window=250\n> > >\n> > > complete seconds RAM\n> > > 10%  47 1GB\n> > > 20%  29 1Gb\n> > > 30%  24 1Gb\n> > > 40%  18 1GB\n> > > 50%  110 1.2GB\n> > > 60%  85 1.4GB\n> > > 70%  195 1.5GB\n> > > 80%  186 2.5GB\n> > > 90%  489 3.8GB\n> > > 95%  800 4.8GB\n> > > I killed it because it started swapping\n> > >\n> > > The mmaps are only about 400MB in this case.\n> > > At the end the git process had 4.4GB of physical RAM allocated.\n> >\n> > That's really bad.\n> >\n> > > Starting from a highly compressed pack greatly aggravates the problem.\n> >\n> > That is really interesting though.\n> >\n> > > Starting with a 2GB pack of the same data my process size only grew to\n> > > 3GB with 2GB of mmaps.\n> >\n> > Which is quite reasonable, even if the same issue might still be there.\n> >\n> > So the problem seems to be related to the pack access code and not the\n> > repack code.  And it must have something to do with the number of deltas\n> > being replayed.  And because the repack is attempting delta compression\n> > roughly from newest to oldest, and because old objects are typically in\n> > a deeper delta chain, then this might explain the logarithmic slowdown.\n> >\n> > So something must be wrong with the delta cache in sha1_file.c somehow.\n\nStaring at the cache code I don't see anything wrong with it.\n\n> I applied the delta accounting patch. It took about 200MB of from the\n> memory use but that doesn't make a dent in 4GB of allocations.\n\nRight.  I didn't expect much from that fix.\n\n\nNicolas\n"},{"id":"62362","messageId":"20071208025600.GA25485@old.davidb.org","threadId":"11182","inReplyTo":"9e4733910712071505y6834f040k37261d65a2d445c4@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"David Brown","fromEmail":"git@davidb.org","sentAt":"2007-12-08T02:56:00Z","receivedAt":"2007-12-08T02:56:00Z","isPatch":false,"sender":{"key":"git@davidb.org","avatar":"https://gravatar.com/avatar/94c86a2938470a74c2eac5e2b69afc0871f79a660295c02219597aba8cb101c1?d=mp&s=160"},"body":"On Fri, Dec 07, 2007 at 06:05:38PM -0500, Jon Smirl wrote:\n>Using this config:\n>[pack]\n>        threads = 4\n>        deltacachesize = 256M\n>        deltacachelimit = 0\n\nJust out of curiousity, does adding\n\n         [pack]\n                 windowmemory = 256M\n\nhelp.  I've found this to grow very large when there are large blobs.\n\nDave\n"},{"id":"62363","messageId":"9e4733910712071929h17a7d88dv37686ec7cd858c63@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712072124160.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-08T03:29:31Z","receivedAt":"2007-12-08T03:29:31Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"The kernel repo has the same problem but not nearly as bad.\n\nStarting from a default pack\n git repack -a -d -f  --depth=1000 --window=1000\nUses 1GB of physical memory\n\nNow do the command again.\n git repack -a -d -f  --depth=1000 --window=1000\nUses 1.3GB of physical memory\n\nI suspect the gcc repo has much longer revision chains than the kernel\none since the kernel repo is only a few years old. The Mozilla repo\ncontained revision chains with over 2,000 revisions. Longer revision\nchains result in longer delta chains.\n\nSo what is allocating the extra memory? Either a function of the\nnumber of entries in the chain, or related to accessing the chain\nsince a chain with more entries will need to be accessed more times.\n\nI have a 168MB kernel pack now after 15 minutes of four cores at 100%.\n\nHere's another observation, the gcc objects are larger. Kernel has\n650K objects in 190MB, gcc has 870K objects in 330MB. Average gcc\nobject is 30% larger. How should the average kernel developer\ninterpret this?\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62364","messageId":"20071208033722.GA27776@old.davidb.org","threadId":"11182","inReplyTo":"9e4733910712071929h17a7d88dv37686ec7cd858c63@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"David Brown","fromEmail":"git@davidb.org","sentAt":"2007-12-08T03:37:22Z","receivedAt":"2007-12-08T03:37:22Z","isPatch":false,"sender":{"key":"git@davidb.org","avatar":"https://gravatar.com/avatar/94c86a2938470a74c2eac5e2b69afc0871f79a660295c02219597aba8cb101c1?d=mp&s=160"},"body":"On Fri, Dec 07, 2007 at 10:29:31PM -0500, Jon Smirl wrote:\n>The kernel repo has the same problem but not nearly as bad.\n>\n>Starting from a default pack\n> git repack -a -d -f  --depth=1000 --window=1000\n>Uses 1GB of physical memory\n>\n>Now do the command again.\n> git repack -a -d -f  --depth=1000 --window=1000\n>Uses 1.3GB of physical memory\n\nWith my repo that contains a bunch of 50MB tarfiles, I've found I must\nspecify --window-memory as well to keep repack from using nearly unbounded\namounts of memory.  Perhaps it is the larger files found in gcc that\nprovokes this.\n\nA window size of 1000 can take a lot of memory if the objects are large.\n\nDave\n"},{"id":"62365","messageId":"1197085456.22471.42.camel@brick","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712072032410.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-08T03:44:16Z","receivedAt":"2007-12-08T03:44:16Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"\nOn Fri, 2007-12-07 at 20:46 -0500, Nicolas Pitre wrote:\n> On Fri, 7 Dec 2007, Jon Smirl wrote:\n> > And the 330MB gcc pack for input\n> >  git repack -a -d -f  --depth=250 --window=250\n> > \n> > complete seconds RAM\n> > 10%  47 1GB\n> > 20%  29 1Gb\n> > 30%  24 1Gb\n> > 40%  18 1GB\n> > 50%  110 1.2GB\n> > 60%  85 1.4GB\n> > 70%  195 1.5GB\n> > 80%  186 2.5GB\n> > 90%  489 3.8GB\n> > 95%  800 4.8GB\n> > I killed it because it started swapping\n> > \n> > The mmaps are only about 400MB in this case.\n> > At the end the git process had 4.4GB of physical RAM allocated.\n> > Starting with a 2GB pack of the same data my process size only grew to\n> > 3GB with 2GB of mmaps.\n> \n> Which is quite reasonable, even if the same issue might still be there.\n> \n> So the problem seems to be related to the pack access code and not the \n> repack code.  And it must have something to do with the number of deltas \n> being replayed.  And because the repack is attempting delta compression \n> roughly from newest to oldest, and because old objects are typically in \n> a deeper delta chain, then this might explain the logarithmic slowdown.\n> \n> So something must be wrong with the delta cache in sha1_file.c somehow.\n\nAll I have is a qualitative observation, but during the process of\ncreating the pack, there was a _huge_ slowdown between 10-15%\n(hundreds/dozens per second to single object per second and a\ncorresponding increase in process size).  Didn't keep any numbers\nat the time, but it was noticable.\n\nI wonder if there are a bunch of huge objects somewhere in gcc's\nhistory?\n\nHarvey\n"},{"id":"62366","messageId":"1197085700.22471.47.camel@brick","threadId":"11182","inReplyTo":"9e4733910712071929h17a7d88dv37686ec7cd858c63@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-08T03:48:20Z","receivedAt":"2007-12-08T03:48:20Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Fri, 2007-12-07 at 22:29 -0500, Jon Smirl wrote:\n> The kernel repo has the same problem but not nearly as bad.\n> \n> Starting from a default pack\n>  git repack -a -d -f  --depth=1000 --window=1000\n> Uses 1GB of physical memory\n> \n> Now do the command again.\n>  git repack -a -d -f  --depth=1000 --window=1000\n> Uses 1.3GB of physical memory\n> \n> I suspect the gcc repo has much longer revision chains than the kernel\n> one since the kernel repo is only a few years old. The Mozilla repo\n> contained revision chains with over 2,000 revisions. Longer revision\n> chains result in longer delta chains.\n\nI sent out a partial delta breakdown for the gcc repo earlier, here's\nthe whole list.\n\nbreakdown of the gcc packfile:\n\nTotal objects\n1017922\n\nChainLength\tObjects\tCumulative\n1:\t103817\t103817\n2:\t67332\t171149\n3:\t57520\t228669\n4:\t52570\t281239\n5:\t43910\t325149\n6:\t37520\t362669\n7:\t35248\t397917\n8:\t29819\t427736\n9:\t27619\t455355\n10:\t22656\t478011\n11:\t21073\t499084\n12:\t18738\t517822\n13:\t16674\t534496\n14:\t14882\t549378\n15:\t14424\t563802\n16:\t12765\t576567\n17:\t11662\t588229\n18:\t11845\t600074\n19:\t11694\t611768\n20:\t9625\t621393\n21:\t9031\t630424\n22:\t8437\t638861\n23:\t8217\t647078\n24:\t7927\t655005\n25:\t7955\t662960\n26:\t7092\t670052\n27:\t7004\t677056\n28:\t6724\t683780\n29:\t6626\t690406\n30:\t5875\t696281\n31:\t5970\t702251\n32:\t5726\t707977\n33:\t6025\t714002\n34:\t5354\t719356\n35:\t6413\t725769\n36:\t4933\t730702\n37:\t4888\t735590\n38:\t4561\t740151\n39:\t4366\t744517\n40:\t4166\t748683\n41:\t4531\t753214\n42:\t4029\t757243\n43:\t3701\t760944\n44:\t3647\t764591\n45:\t3553\t768144\n46:\t3509\t771653\n47:\t3473\t775126\n48:\t3442\t778568\n49:\t3379\t781947\n50:\t3395\t785342\n51:\t3315\t788657\n52:\t3168\t791825\n53:\t3345\t795170\n54:\t3166\t798336\n55:\t3237\t801573\n56:\t2795\t804368\n57:\t2768\t807136\n58:\t2666\t809802\n59:\t2723\t812525\n60:\t2547\t815072\n61:\t2565\t817637\n62:\t2622\t820259\n63:\t2521\t822780\n64:\t2492\t825272\n65:\t2529\t827801\n66:\t2566\t830367\n67:\t2685\t833052\n68:\t2458\t835510\n69:\t2457\t837967\n70:\t2440\t840407\n71:\t2410\t842817\n72:\t2337\t845154\n73:\t2301\t847455\n74:\t2201\t849656\n75:\t2127\t851783\n76:\t2256\t854039\n77:\t2038\t856077\n78:\t1925\t858002\n79:\t1965\t859967\n80:\t1929\t861896\n81:\t1890\t863786\n82:\t1873\t865659\n83:\t1964\t867623\n84:\t1898\t869521\n85:\t1839\t871360\n86:\t1933\t873293\n87:\t1876\t875169\n88:\t1851\t877020\n89:\t1789\t878809\n90:\t1790\t880599\n91:\t1804\t882403\n92:\t1696\t884099\n93:\t1863\t885962\n94:\t1889\t887851\n95:\t1766\t889617\n96:\t1731\t891348\n97:\t1775\t893123\n98:\t1750\t894873\n99:\t1767\t896640\n100:\t1644\t898284\n101:\t1642\t899926\n102:\t1489\t901415\n103:\t1532\t902947\n104:\t1564\t904511\n105:\t1477\t905988\n106:\t1461\t907449\n107:\t1383\t908832\n108:\t1422\t910254\n109:\t1316\t911570\n110:\t1480\t913050\n111:\t1329\t914379\n112:\t1375\t915754\n113:\t1292\t917046\n114:\t1224\t918270\n115:\t1123\t919393\n116:\t1216\t920609\n117:\t1252\t921861\n118:\t1252\t923113\n119:\t1346\t924459\n120:\t1320\t925779\n121:\t1277\t927056\n122:\t1234\t928290\n123:\t1200\t929490\n124:\t1255\t930745\n125:\t1206\t931951\n126:\t1155\t933106\n127:\t1246\t934352\n128:\t1226\t935578\n129:\t1194\t936772\n130:\t1268\t938040\n131:\t1334\t939374\n132:\t1146\t940520\n133:\t1220\t941740\n134:\t1055\t942795\n135:\t1110\t943905\n136:\t1095\t945000\n137:\t1294\t946294\n138:\t1204\t947498\n139:\t1218\t948716\n140:\t1101\t949817\n141:\t993\t950810\n142:\t975\t951785\n143:\t1014\t952799\n144:\t968\t953767\n145:\t957\t954724\n146:\t1069\t955793\n147:\t996\t956789\n148:\t967\t957756\n149:\t964\t958720\n150:\t954\t959674\n151:\t949\t960623\n152:\t1001\t961624\n153:\t1042\t962666\n154:\t1057\t963723\n155:\t948\t964671\n156:\t966\t965637\n157:\t833\t966470\n158:\t959\t967429\n159:\t907\t968336\n160:\t854\t969190\n161:\t847\t970037\n162:\t836\t970873\n163:\t769\t971642\n164:\t747\t972389\n165:\t755\t973144\n166:\t707\t973851\n167:\t774\t974625\n168:\t777\t975402\n169:\t783\t976185\n170:\t707\t976892\n171:\t738\t977630\n172:\t775\t978405\n173:\t781\t979186\n174:\t698\t979884\n175:\t801\t980685\n176:\t712\t981397\n177:\t679\t982076\n178:\t775\t982851\n179:\t696\t983547\n180:\t760\t984307\n181:\t740\t985047\n182:\t752\t985799\n183:\t704\t986503\n184:\t683\t987186\n185:\t690\t987876\n186:\t741\t988617\n187:\t642\t989259\n188:\t672\t989931\n189:\t679\t990610\n190:\t691\t991301\n191:\t648\t991949\n192:\t703\t992652\n193:\t675\t993327\n194:\t687\t994014\n195:\t625\t994639\n196:\t607\t995246\n197:\t583\t995829\n198:\t632\t996461\n199:\t540\t997001\n200:\t652\t997653\n201:\t600\t998253\n202:\t628\t998881\n203:\t624\t999505\n204:\t582\t1000087\n205:\t548\t1000635\n206:\t520\t1001155\n207:\t648\t1001803\n208:\t556\t1002359\n209:\t563\t1002922\n210:\t508\t1003430\n211:\t570\t1004000\n212:\t530\t1004530\n213:\t575\t1005105\n214:\t527\t1005632\n215:\t521\t1006153\n216:\t515\t1006668\n217:\t513\t1007181\n218:\t460\t1007641\n219:\t491\t1008132\n220:\t474\t1008606\n221:\t471\t1009077\n222:\t482\t1009559\n223:\t485\t1010044\n224:\t439\t1010483\n225:\t385\t1010868\n226:\t385\t1011253\n227:\t403\t1011656\n228:\t380\t1012036\n229:\t376\t1012412\n230:\t377\t1012789\n231:\t415\t1013204\n232:\t394\t1013598\n233:\t362\t1013960\n234:\t334\t1014294\n235:\t366\t1014660\n236:\t317\t1014977\n237:\t362\t1015339\n238:\t343\t1015682\n239:\t392\t1016074\n240:\t317\t1016391\n241:\t305\t1016696\n242:\t319\t1017015\n243:\t276\t1017291\n244:\t247\t1017538\n245:\t179\t1017717\n246:\t111\t1017828\n247:\t61\t1017889\n248:\t27\t1017916\n249:\t6\t1017922\n\nHarvey\n"},{"id":"62368","messageId":"9e4733910712072022na3369caob48d4b26a56224ea@mail.gmail.com","threadId":"11182","inReplyTo":"20071208033722.GA27776@old.davidb.org","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-08T04:22:02Z","receivedAt":"2007-12-08T04:22:02Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/7/07, David Brown <git@davidb.org> wrote:\n> On Fri, Dec 07, 2007 at 10:29:31PM -0500, Jon Smirl wrote:\n> >The kernel repo has the same problem but not nearly as bad.\n> >\n> >Starting from a default pack\n> > git repack -a -d -f  --depth=1000 --window=1000\n> >Uses 1GB of physical memory\n> >\n> >Now do the command again.\n> > git repack -a -d -f  --depth=1000 --window=1000\n> >Uses 1.3GB of physical memory\n>\n> With my repo that contains a bunch of 50MB tarfiles, I've found I must\n> specify --window-memory as well to keep repack from using nearly unbounded\n> amounts of memory.  Perhaps it is the larger files found in gcc that\n> provokes this.\n>\n> A window size of 1000 can take a lot of memory if the objects are large.\n\nThis is a partial solution to the problem. Adding window size =256M\ntook memory consumption down from 4.8GB to 2.8GB. It took an hour to\nrun the test.\n\nIt not the complete solution since my git process is still using 2.4GB\nphysical memory. I also still experiencing a lot of slow down in the\nlast 10%.\n\nDoes the gcc repo contain some giant objects? Why wasn't the memory\nfreed after their chain was processed?\n\nMost of the last 10% is being done on a single CPU. There must be a\nchain of giant objects that is unbalancing everything.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62369","messageId":"alpine.LFD.0.99999.0712072328420.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712072022na3369caob48d4b26a56224ea@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-08T04:30:16Z","receivedAt":"2007-12-08T04:30:16Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Dec 2007, Jon Smirl wrote:\n\n> Does the gcc repo contain some giant objects? Why wasn't the memory\n> freed after their chain was processed?\n\nIt should be.\n\n> Most of the last 10% is being done on a single CPU. There must be a\n> chain of giant objects that is unbalancing everything.\n\nI'm about to send a patch to fix the thread balancing for real this \ntime.\n\n\nNicolas\n"},{"id":"62373","messageId":"9e4733910712072101k4583c0afsea368253fe1cf706@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712072328420.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-08T05:01:13Z","receivedAt":"2007-12-08T05:01:13Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/7/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 7 Dec 2007, Jon Smirl wrote:\n>\n> > Does the gcc repo contain some giant objects? Why wasn't the memory\n> > freed after their chain was processed?\n>\n> It should be.\n>\n> > Most of the last 10% is being done on a single CPU. There must be a\n> > chain of giant objects that is unbalancing everything.\n>\n> I'm about to send a patch to fix the thread balancing for real this\n> time.\n\nSomething is really broken in the last 5% of that repo. I have been\nprocessing at 97% for 30 minutes without moving to 98%.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62376","messageId":"alpine.LFD.0.99999.0712080004490.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712072101k4583c0afsea368253fe1cf706@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-08T05:12:53Z","receivedAt":"2007-12-08T05:12:53Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 8 Dec 2007, Jon Smirl wrote:\n\n> On 12/7/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Fri, 7 Dec 2007, Jon Smirl wrote:\n> >\n> > > Does the gcc repo contain some giant objects? Why wasn't the memory\n> > > freed after their chain was processed?\n> >\n> > It should be.\n> >\n> > > Most of the last 10% is being done on a single CPU. There must be a\n> > > chain of giant objects that is unbalancing everything.\n> >\n> > I'm about to send a patch to fix the thread balancing for real this\n> > time.\n> \n> Something is really broken in the last 5% of that repo. I have been\n> processing at 97% for 30 minutes without moving to 98%.\n\nThis is a clear sign of a problem, indeed.\n\nI'll be away for the weekend, so here's a few things to try out if you \nfeel like it:\n\n1) Make sure the problem occurs with the thread code disabled.  That \n   would eliminate one variable, and will help for #2.\n\n2) Try bissecting the issue.  If you can find an old Git version where \n   the issue doesn't appear then simply run \"git bissect\" to find the \n   exact commit causing the problem.  Best with a repo that doesn't take\n   ages to repack.\n\n3) Compile Git against the dmalloc library in order to identify where\n   the huge memory leak is happening.\n\n\nNicolas\n"},{"id":"62434","messageId":"7vodd0vnhv.fsf@gitster.siamese.dyndns.org","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712072032410.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-08T22:18:52Z","receivedAt":"2007-12-08T22:18:52Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Fri, 7 Dec 2007, Jon Smirl wrote:\n>\n>> Starting with a 2GB pack of the same data my process size only grew to\n>> 3GB with 2GB of mmaps.\n>\n> Which is quite reasonable, even if the same issue might still be there.\n>\n> So the problem seems to be related to the pack access code and not the \n> repack code.  And it must have something to do with the number of deltas \n> being replayed.  And because the repack is attempting delta compression \n> roughly from newest to oldest, and because old objects are typically in \n> a deeper delta chain, then this might explain the logarithmic slowdown.\n>\n> So something must be wrong with the delta cache in sha1_file.c somehow.\n\nI was reaching the same conclusion but haven't managed to spot anything\nblatantly wrong in that area.  Will need to dig more.\n"},{"id":"62464","messageId":"7vprxgs36w.fsf@gitster.siamese.dyndns.org","threadId":"11182","inReplyTo":"7vodd0vnhv.fsf@gitster.siamese.dyndns.org","subject":"Re: Something is broken in repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-09T08:05:43Z","receivedAt":"2007-12-09T08:05:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Nicolas Pitre <nico@cam.org> writes:\n>\n>> On Fri, 7 Dec 2007, Jon Smirl wrote:\n>>\n>>> Starting with a 2GB pack of the same data my process size only grew to\n>>> 3GB with 2GB of mmaps.\n>>\n>> Which is quite reasonable, even if the same issue might still be there.\n>>\n>> So the problem seems to be related to the pack access code and not the \n>> repack code.  And it must have something to do with the number of deltas \n>> being replayed.  And because the repack is attempting delta compression \n>> roughly from newest to oldest, and because old objects are typically in \n>> a deeper delta chain, then this might explain the logarithmic slowdown.\n>>\n>> So something must be wrong with the delta cache in sha1_file.c somehow.\n>\n> I was reaching the same conclusion but haven't managed to spot anything\n> blatantly wrong in that area.  Will need to dig more.\n\nDoes this problem have correlation with the use of threads?  Do you see\nthe same bloat with or without THREADED_DELTA_SEARCH defined?\n"},{"id":"62486","messageId":"9e4733910712090719g713972a9p1c18bc149dc0237c@mail.gmail.com","threadId":"11182","inReplyTo":"7vprxgs36w.fsf@gitster.siamese.dyndns.org","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-09T15:19:23Z","receivedAt":"2007-12-09T15:19:23Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/9/07, Junio C Hamano <gitster@pobox.com> wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > Nicolas Pitre <nico@cam.org> writes:\n> >\n> >> On Fri, 7 Dec 2007, Jon Smirl wrote:\n> >>\n> >>> Starting with a 2GB pack of the same data my process size only grew to\n> >>> 3GB with 2GB of mmaps.\n> >>\n> >> Which is quite reasonable, even if the same issue might still be there.\n> >>\n> >> So the problem seems to be related to the pack access code and not the\n> >> repack code.  And it must have something to do with the number of deltas\n> >> being replayed.  And because the repack is attempting delta compression\n> >> roughly from newest to oldest, and because old objects are typically in\n> >> a deeper delta chain, then this might explain the logarithmic slowdown.\n> >>\n> >> So something must be wrong with the delta cache in sha1_file.c somehow.\n> >\n> > I was reaching the same conclusion but haven't managed to spot anything\n> > blatantly wrong in that area.  Will need to dig more.\n>\n> Does this problem have correlation with the use of threads?  Do you see\n> the same bloat with or without THREADED_DELTA_SEARCH defined?\n>\n\nI just started a non-threaded one. It will be four or five hours\nbefore it finishes.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62507","messageId":"9e4733910712091025s27c3d698yba78eed4306cd3ec@mail.gmail.com","threadId":"11182","inReplyTo":"7vprxgs36w.fsf@gitster.siamese.dyndns.org","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-09T18:25:04Z","receivedAt":"2007-12-09T18:25:04Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/9/07, Junio C Hamano <gitster@pobox.com> wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n> > Nicolas Pitre <nico@cam.org> writes:\n> >\n> >> On Fri, 7 Dec 2007, Jon Smirl wrote:\n> >>\n> >>> Starting with a 2GB pack of the same data my process size only grew to\n> >>> 3GB with 2GB of mmaps.\n> >>\n> >> Which is quite reasonable, even if the same issue might still be there.\n> >>\n> >> So the problem seems to be related to the pack access code and not the\n> >> repack code.  And it must have something to do with the number of deltas\n> >> being replayed.  And because the repack is attempting delta compression\n> >> roughly from newest to oldest, and because old objects are typically in\n> >> a deeper delta chain, then this might explain the logarithmic slowdown.\n> >>\n> >> So something must be wrong with the delta cache in sha1_file.c somehow.\n> >\n> > I was reaching the same conclusion but haven't managed to spot anything\n> > blatantly wrong in that area.  Will need to dig more.\n>\n> Does this problem have correlation with the use of threads?  Do you see\n> the same bloat with or without THREADED_DELTA_SEARCH defined?\n>\n\nSomething else seems to be wrong.\n\nWith threading turned off,  5000 CPU seconds and 13% done.\nWith threading turned on, threads = 1, 5000 CPU seconds, 13%\nWith threading turned on, threads = 2, 180 CPU seconds, 13%\nWith threading turned on, threads = 4, 150 CPU seconds, 13%\n\nThis can't be right, four cores are not 40x one core. So maybe the\nobserved logarithmic slow down is because the percent complete is\nbeing reported wrong in the threaded case. If that's the case we may\nbe looking in the wrong place for problems.\n\nThe times are only approximate, I'm using the CPU for other things.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62528","messageId":"alpine.LFD.0.99999.0712092002170.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712091025s27c3d698yba78eed4306cd3ec@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-10T01:07:49Z","receivedAt":"2007-12-10T01:07:49Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 9 Dec 2007, Jon Smirl wrote:\n\n> On 12/9/07, Junio C Hamano <gitster@pobox.com> wrote:\n> > Junio C Hamano <gitster@pobox.com> writes:\n> >\n> > > Nicolas Pitre <nico@cam.org> writes:\n> > >\n> > >> On Fri, 7 Dec 2007, Jon Smirl wrote:\n> > >>\n> > >>> Starting with a 2GB pack of the same data my process size only grew to\n> > >>> 3GB with 2GB of mmaps.\n> > >>\n> > >> Which is quite reasonable, even if the same issue might still be there.\n> > >>\n> > >> So the problem seems to be related to the pack access code and not the\n> > >> repack code.  And it must have something to do with the number of deltas\n> > >> being replayed.  And because the repack is attempting delta compression\n> > >> roughly from newest to oldest, and because old objects are typically in\n> > >> a deeper delta chain, then this might explain the logarithmic slowdown.\n> > >>\n> > >> So something must be wrong with the delta cache in sha1_file.c somehow.\n> > >\n> > > I was reaching the same conclusion but haven't managed to spot anything\n> > > blatantly wrong in that area.  Will need to dig more.\n> >\n> > Does this problem have correlation with the use of threads?  Do you see\n> > the same bloat with or without THREADED_DELTA_SEARCH defined?\n> >\n> \n> Something else seems to be wrong.\n> \n> With threading turned off,  5000 CPU seconds and 13% done.\n> With threading turned on, threads = 1, 5000 CPU seconds, 13%\n> With threading turned on, threads = 2, 180 CPU seconds, 13%\n> With threading turned on, threads = 4, 150 CPU seconds, 13%\n> \n> This can't be right, four cores are not 40x one core.\n\nIt may be right.  The object list to apply delta compression on doesn't \nnecessarily require a uniform amount of cycles throughout.  When using \nmultiple threads, the list is broken in parts for each thread, and later \nparts might end up being simply much easier to process, therefore \nchanging the percentage figure.\n\n> So maybe the observed logarithmic slow down is because the percent \n> complete is being reported wrong in the threaded case. If that's the \n> case we may be looking in the wrong place for problems.\n\nI really doubt it.\n\n\nNicolas\n"},{"id":"62531","messageId":"alpine.LFD.0.99999.0712092144220.555@xanadu.home","threadId":"11182","inReplyTo":"7vodd0vnhv.fsf@gitster.siamese.dyndns.org","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-10T02:49:32Z","receivedAt":"2007-12-10T02:49:32Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 8 Dec 2007, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > On Fri, 7 Dec 2007, Jon Smirl wrote:\n> >\n> >> Starting with a 2GB pack of the same data my process size only grew to\n> >> 3GB with 2GB of mmaps.\n> >\n> > Which is quite reasonable, even if the same issue might still be there.\n> >\n> > So the problem seems to be related to the pack access code and not the \n> > repack code.  And it must have something to do with the number of deltas \n> > being replayed.  And because the repack is attempting delta compression \n> > roughly from newest to oldest, and because old objects are typically in \n> > a deeper delta chain, then this might explain the logarithmic slowdown.\n> >\n> > So something must be wrong with the delta cache in sha1_file.c somehow.\n> \n> I was reaching the same conclusion but haven't managed to spot anything\n> blatantly wrong in that area.  Will need to dig more.\n\nI didn't find anything wrong there either. I'll have to run some more \ngcc repacking tests myself, despite not having a blazingly fast machine \nmaking for rather long turnarounds.\n\n\nNicolas\n"},{"id":"62610","messageId":"alpine.LFD.0.99999.0712101434560.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712071505y6834f040k37261d65a2d445c4@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-10T19:56:21Z","receivedAt":"2007-12-10T19:56:21Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 7 Dec 2007, Jon Smirl wrote:\n\n> Using this config:\n> [pack]\n>         threads = 4\n>         deltacachesize = 256M\n>         deltacachelimit = 0\n> \n> And the 330MB gcc pack for input\n>  git repack -a -d -f  --depth=250 --window=250\n> \n> complete seconds RAM\n> 10%  47 1GB\n> 20%  29 1Gb\n> 30%  24 1Gb\n> 40%  18 1GB\n> 50%  110 1.2GB\n> 60%  85 1.4GB\n> 70%  195 1.5GB\n> 80%  186 2.5GB\n> 90%  489 3.8GB\n> 95%  800 4.8GB\n> I killed it because it started swapping\n> \n> The mmaps are only about 400MB in this case.\n> At the end the git process had 4.4GB of physical RAM allocated.\n> \n> Starting from a highly compressed pack greatly aggravates the problem.\n> Starting with a 2GB pack of the same data my process size only grew to\n> 3GB with 2GB of mmaps.\n\nYou said having reproduced the issue, albeit not as severe, with the \nLinux kernel repo.  I did just that:\n\n# to get the default pack:\n$ git repack -a -f -d\n\n# first measurement with a repack from a default pack\n$ /usr/bin/time git repack -a -f --window=256 --depth=256\n2572.17user 5.87system 22:46.80elapsed 188%CPU (0avgtext+0avgdata 0maxresident)k\n15720inputs+356640outputs (71major+264376minor)pagefaults 0swaps\n\n# do it again to start from a highly packed pack\n$ /usr/bin/time git repack -a -f --window=256 --depth=256\n2573.53user 5.62system 22:45.60elapsed 188%CPU (0avgtext+0avgdata 0maxresident)k\n29176inputs+356664outputs (210major+274887minor)pagefaults 0swaps\n\nThis is with pack.threads=2 on a P4 with HT, and I'm using the machine \nfor other tasks as well, but all measured time is sensibly the same for \nboth cases.  Virtual memory allocation never reached 700MB in both cases \neither.\n\n\nNicolas\n"},{"id":"62611","messageId":"9e4733910712101205q218152a2td14a8931e63d2610@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712101434560.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-10T20:05:53Z","receivedAt":"2007-12-10T20:05:53Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/10/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Fri, 7 Dec 2007, Jon Smirl wrote:\n>\n> > Using this config:\n> > [pack]\n> >         threads = 4\n> >         deltacachesize = 256M\n> >         deltacachelimit = 0\n> >\n> > And the 330MB gcc pack for input\n> >  git repack -a -d -f  --depth=250 --window=250\n> >\n> > complete seconds RAM\n> > 10%  47 1GB\n> > 20%  29 1Gb\n> > 30%  24 1Gb\n> > 40%  18 1GB\n> > 50%  110 1.2GB\n> > 60%  85 1.4GB\n> > 70%  195 1.5GB\n> > 80%  186 2.5GB\n> > 90%  489 3.8GB\n> > 95%  800 4.8GB\n> > I killed it because it started swapping\n> >\n> > The mmaps are only about 400MB in this case.\n> > At the end the git process had 4.4GB of physical RAM allocated.\n> >\n> > Starting from a highly compressed pack greatly aggravates the problem.\n> > Starting with a 2GB pack of the same data my process size only grew to\n> > 3GB with 2GB of mmaps.\n>\n> You said having reproduced the issue, albeit not as severe, with the\n> Linux kernel repo.  I did just that:\n>\n> # to get the default pack:\n> $ git repack -a -f -d\n>\n> # first measurement with a repack from a default pack\n> $ /usr/bin/time git repack -a -f --window=256 --depth=256\n> 2572.17user 5.87system 22:46.80elapsed 188%CPU (0avgtext+0avgdata 0maxresident)k\n> 15720inputs+356640outputs (71major+264376minor)pagefaults 0swaps\n>\n> # do it again to start from a highly packed pack\n> $ /usr/bin/time git repack -a -f --window=256 --depth=256\n> 2573.53user 5.62system 22:45.60elapsed 188%CPU (0avgtext+0avgdata 0maxresident)k\n> 29176inputs+356664outputs (210major+274887minor)pagefaults 0swaps\n>\n> This is with pack.threads=2 on a P4 with HT, and I'm using the machine\n> for other tasks as well, but all measured time is sensibly the same for\n> both cases.  Virtual memory allocation never reached 700MB in both cases\n> either.\n>\n\nThis is the mail about the kernel pack, the one you quoted is a gcc run.\n\nThe kernel repo has the same problem but not nearly as bad.\n\nStarting from a default pack\n git repack -a -d -f  --depth=1000 --window=1000\nUses 1GB of physical memory\n\nNow do the command again.\n git repack -a -d -f  --depth=1000 --window=1000\nUses 1.3GB of physical memory\n\nI suspect the gcc repo has much longer revision chains than the kernel\none since the kernel repo is only a few years old. The Mozilla repo\ncontained revision chains with over 2,000 revisions. Longer revision\nchains result in longer delta chains.\n\nSo what is allocating the extra memory? Either a function of the\nnumber of entries in the chain, or related to accessing the chain\nsince a chain with more entries will need to be accessed more times.\n\nI have a 168MB kernel pack now after 15 minutes of four cores at 100%.\n\nHere's another observation, the gcc objects are larger. Kernel has\n650K objects in 190MB, gcc has 870K objects in 330MB. Average gcc\nobject is 30% larger. How should the average kernel developer\ninterpret this?\n\n\n\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62613","messageId":"118833cc0712101216x989e720pe190c60025409bd2@mail.gmail.com","threadId":"11182","inReplyTo":"9e4733910712101205q218152a2td14a8931e63d2610@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2007-12-10T20:16:50Z","receivedAt":"2007-12-10T20:16:50Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"> Here's another observation, the gcc objects are larger. Kernel has\n> 650K objects in 190MB, gcc has 870K objects in 330MB. Average gcc\n> object is 30% larger. How should the average kernel developer\n> interpret this?\n\nCould this be explained by the ChangeLog file?  It's large; it has tons of\nrevisions; it is a prime candidate for delta compression.\n\nMorten\n"},{"id":"62637","messageId":"9e4733910712101825l33cdc2c0mca2ddbfd5afdb298@mail.gmail.com","threadId":"11182","inReplyTo":"9e4733910712071505y6834f040k37261d65a2d445c4@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T02:25:26Z","receivedAt":"2007-12-11T02:25:26Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"New run using same configuration. With the addition of the more\nefficient load balancing patches and delta cache accounting.\n\nSeconds are wall clock time. They are lower since the patch made\nthreading better at using all four cores. I am stuck at 380-390% CPU\nutilization for the git process.\n\ncomplete seconds RAM\n10%   60    900M (includes counting)\n20%   15    900M\n30%   15    900M\n40%   50    1.2G\n50%   80    1.3G\n60%   70    1.7G\n70%   140  1.8G\n80%   180  2.0G\n90%   280  2.2G\n95%   530  2.8G - 1,420 total to here, previous was 1,983\n100% 1390 2.85G\nDuring the writing phase RAM fell to 1.6G\nWhat is being freed in the writing phase??\n\nI have no explanation for the change in RAM usage. Two guesses come to\nmind. Memory fragmentation. Or the change in the way the work was\nsplit up altered RAM usage.\n\nTotal CPU time was 195 minutes in 70 minutes clock time. About 70%\nefficient. During the compress phase all four cores were active until\nthe last 90 seconds. Writing the objects took over 23 minutes CPU\nbound on one core.\n\nNew pack file is: 270,594,853\nOld one was: 344,543,752\nIt still has 828,660 objects\n\n\nOn 12/7/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> Using this config:\n> [pack]\n>         threads = 4\n>         deltacachesize = 256M\n>         deltacachelimit = 0\n>\n> And the 330MB gcc pack for input\n>  git repack -a -d -f  --depth=250 --window=250\n>\n> complete seconds RAM\n> 10%  47 1GB\n> 20%  29 1Gb\n> 30%  24 1Gb\n> 40%  18 1GB\n> 50%  110 1.2GB\n> 60%  85 1.4GB\n> 70%  195 1.5GB\n> 80%  186 2.5GB\n> 90%  489 3.8GB\n> 95%  800 4.8GB\n> I killed it because it started swapping\n>\n> The mmaps are only about 400MB in this case.\n> At the end the git process had 4.4GB of physical RAM allocated.\n>\n> Starting from a highly compressed pack greatly aggravates the problem.\n> Starting with a 2GB pack of the same data my process size only grew to\n> 3GB with 2GB of mmaps.\n>\n> --\n> Jon Smirl\n> jonsmirl@gmail.com\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62640","messageId":"7vlk82hrdt.fsf@gitster.siamese.dyndns.org","threadId":"11182","inReplyTo":"9e4733910712101825l33cdc2c0mca2ddbfd5afdb298@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-11T02:55:26Z","receivedAt":"2007-12-11T02:55:26Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Jon Smirl\" <jonsmirl@gmail.com> writes:\n\n> 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> 100% 1390 2.85G\n> During the writing phase RAM fell to 1.6G\n> What is being freed in the writing phase??\n\nentry->delta_data is the only thing I can think of that are freed\nin the function that have been allocated much earlier before entering\nthe function.\n"},{"id":"62642","messageId":"alpine.LFD.0.99999.0712102225240.555@xanadu.home","threadId":"11182","inReplyTo":"7vlk82hrdt.fsf@gitster.siamese.dyndns.org","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T03:27:03Z","receivedAt":"2007-12-11T03:27:03Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 10 Dec 2007, Junio C Hamano wrote:\n\n> \"Jon Smirl\" <jonsmirl@gmail.com> writes:\n> \n> > 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> > 100% 1390 2.85G\n> > During the writing phase RAM fell to 1.6G\n> > What is being freed in the writing phase??\n> \n> entry->delta_data is the only thing I can think of that are freed\n> in the function that have been allocated much earlier before entering\n> the function.\n\nYet all ->delta-data instances are limited to 256MB according to Jon's \nconfig.\n\n\nNicolas\n\n> \n\n\nNicolas\n"},{"id":"62643","messageId":"alpine.LFD.0.99999.0712102231570.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712101825l33cdc2c0mca2ddbfd5afdb298@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T03:49:00Z","receivedAt":"2007-12-11T03:49:00Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 10 Dec 2007, Jon Smirl wrote:\n\n> New run using same configuration. With the addition of the more\n> efficient load balancing patches and delta cache accounting.\n> \n> Seconds are wall clock time. They are lower since the patch made\n> threading better at using all four cores. I am stuck at 380-390% CPU\n> utilization for the git process.\n> \n> complete seconds RAM\n> 10%   60    900M (includes counting)\n> 20%   15    900M\n> 30%   15    900M\n> 40%   50    1.2G\n> 50%   80    1.3G\n> 60%   70    1.7G\n> 70%   140  1.8G\n> 80%   180  2.0G\n> 90%   280  2.2G\n> 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> 100% 1390 2.85G\n> During the writing phase RAM fell to 1.6G\n> What is being freed in the writing phase??\n\nThe cached delta results, but you put a cap of 256MB for them.\n\nCould you try again with that cache disabled entirely, with \npack.deltacachesize = 1 (don't use 0 as that means unbounded).\n\nAnd then, while still keeping the delta cache disabled, could you try \nwith pack.threads = 2, and pack.threads = 1 ?\n\nI'm sorry to ask you to do this but I don't have enough ram to even \ncomplete a repack with threads=2 so I'm reattempting single threaded at \nthe moment.  But I really wonder if the threading has such an effect on \nmemory usage.\n\n\n\n> \n> I have no explanation for the change in RAM usage. Two guesses come to\n> mind. Memory fragmentation. Or the change in the way the work was\n> split up altered RAM usage.\n> \n> Total CPU time was 195 minutes in 70 minutes clock time. About 70%\n> efficient. During the compress phase all four cores were active until\n> the last 90 seconds. Writing the objects took over 23 minutes CPU\n> bound on one core.\n> \n> New pack file is: 270,594,853\n> Old one was: 344,543,752\n> It still has 828,660 objects\n\nYou mean the pack for the gcc repo is now less than 300MB?  Wow.\n\n\nNicolas\n"},{"id":"62652","messageId":"9e4733910712102125w56c70c0cxb8b00a060b62077@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712102231570.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T05:25:55Z","receivedAt":"2007-12-11T05:25:55Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/10/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Mon, 10 Dec 2007, Jon Smirl wrote:\n>\n> > New run using same configuration. With the addition of the more\n> > efficient load balancing patches and delta cache accounting.\n> >\n> > Seconds are wall clock time. They are lower since the patch made\n> > threading better at using all four cores. I am stuck at 380-390% CPU\n> > utilization for the git process.\n> >\n> > complete seconds RAM\n> > 10%   60    900M (includes counting)\n> > 20%   15    900M\n> > 30%   15    900M\n> > 40%   50    1.2G\n> > 50%   80    1.3G\n> > 60%   70    1.7G\n> > 70%   140  1.8G\n> > 80%   180  2.0G\n> > 90%   280  2.2G\n> > 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> > 100% 1390 2.85G\n> > During the writing phase RAM fell to 1.6G\n> > What is being freed in the writing phase??\n>\n> The cached delta results, but you put a cap of 256MB for them.\n>\n> Could you try again with that cache disabled entirely, with\n> pack.deltacachesize = 1 (don't use 0 as that means unbounded).\n>\n> And then, while still keeping the delta cache disabled, could you try\n> with pack.threads = 2, and pack.threads = 1 ?\n>\n> I'm sorry to ask you to do this but I don't have enough ram to even\n> complete a repack with threads=2 so I'm reattempting single threaded at\n> the moment.  But I really wonder if the threading has such an effect on\n> memory usage.\n\nI already have a threads = 1 running with this config. Binary and\nconfig were same from threads=4 run.\n\n10% 28min 950M\n40% 135min 950M\n50% 157min 900M\n60% 160min 830M\n100% 170min 830M\n\nSomething is hurting bad with threads. 170 CPU minutes with one\nthread, versus 195 CPU minutes with four threads.\n\nIs there a different memory allocator that can be used when\nmultithreaded on gcc? This whole problem may be coming from the memory\nallocation function. git is hardly interacting at all on the thread\nlevel so it's likely a problem in the C run-time.\n\n[core]\n        repositoryformatversion = 0\n        filemode = true\n        bare = false\n        logallrefupdates = true\n[pack]\n        threads = 1\n        deltacachesize = 256M\n        windowmemory = 256M\n        deltacachelimit = 0\n[remote \"origin\"]\n        url = git://git.infradead.org/gcc.git\n        fetch = +refs/heads/*:refs/remotes/origin/*\n[branch \"trunk\"]\n        remote = origin\n        merge = refs/heads/trunk\n\n\n\n\n>\n>\n>\n> >\n> > I have no explanation for the change in RAM usage. Two guesses come to\n> > mind. Memory fragmentation. Or the change in the way the work was\n> > split up altered RAM usage.\n> >\n> > Total CPU time was 195 minutes in 70 minutes clock time. About 70%\n> > efficient. During the compress phase all four cores were active until\n> > the last 90 seconds. Writing the objects took over 23 minutes CPU\n> > bound on one core.\n> >\n> > New pack file is: 270,594,853\n> > Old one was: 344,543,752\n> > It still has 828,660 objects\n>\n> You mean the pack for the gcc repo is now less than 300MB?  Wow.\n>\n>\n> Nicolas\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62654","messageId":"9e4733910712102129v140c2affqf2e73e75855b61ea@mail.gmail.com","threadId":"11182","inReplyTo":"9e4733910712102125w56c70c0cxb8b00a060b62077@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T05:29:22Z","receivedAt":"2007-12-11T05:29:22Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I added the gcc people to the CC, it's their repository. Maybe they\ncan help up sort this out.\n\nOn 12/11/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> On 12/10/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Mon, 10 Dec 2007, Jon Smirl wrote:\n> >\n> > > New run using same configuration. With the addition of the more\n> > > efficient load balancing patches and delta cache accounting.\n> > >\n> > > Seconds are wall clock time. They are lower since the patch made\n> > > threading better at using all four cores. I am stuck at 380-390% CPU\n> > > utilization for the git process.\n> > >\n> > > complete seconds RAM\n> > > 10%   60    900M (includes counting)\n> > > 20%   15    900M\n> > > 30%   15    900M\n> > > 40%   50    1.2G\n> > > 50%   80    1.3G\n> > > 60%   70    1.7G\n> > > 70%   140  1.8G\n> > > 80%   180  2.0G\n> > > 90%   280  2.2G\n> > > 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> > > 100% 1390 2.85G\n> > > During the writing phase RAM fell to 1.6G\n> > > What is being freed in the writing phase??\n> >\n> > The cached delta results, but you put a cap of 256MB for them.\n> >\n> > Could you try again with that cache disabled entirely, with\n> > pack.deltacachesize = 1 (don't use 0 as that means unbounded).\n> >\n> > And then, while still keeping the delta cache disabled, could you try\n> > with pack.threads = 2, and pack.threads = 1 ?\n> >\n> > I'm sorry to ask you to do this but I don't have enough ram to even\n> > complete a repack with threads=2 so I'm reattempting single threaded at\n> > the moment.  But I really wonder if the threading has such an effect on\n> > memory usage.\n>\n> I already have a threads = 1 running with this config. Binary and\n> config were same from threads=4 run.\n>\n> 10% 28min 950M\n> 40% 135min 950M\n> 50% 157min 900M\n> 60% 160min 830M\n> 100% 170min 830M\n>\n> Something is hurting bad with threads. 170 CPU minutes with one\n> thread, versus 195 CPU minutes with four threads.\n>\n> Is there a different memory allocator that can be used when\n> multithreaded on gcc? This whole problem may be coming from the memory\n> allocation function. git is hardly interacting at all on the thread\n> level so it's likely a problem in the C run-time.\n>\n> [core]\n>         repositoryformatversion = 0\n>         filemode = true\n>         bare = false\n>         logallrefupdates = true\n> [pack]\n>         threads = 1\n>         deltacachesize = 256M\n>         windowmemory = 256M\n>         deltacachelimit = 0\n> [remote \"origin\"]\n>         url = git://git.infradead.org/gcc.git\n>         fetch = +refs/heads/*:refs/remotes/origin/*\n> [branch \"trunk\"]\n>         remote = origin\n>         merge = refs/heads/trunk\n>\n>\n>\n>\n> >\n> >\n> >\n> > >\n> > > I have no explanation for the change in RAM usage. Two guesses come to\n> > > mind. Memory fragmentation. Or the change in the way the work was\n> > > split up altered RAM usage.\n> > >\n> > > Total CPU time was 195 minutes in 70 minutes clock time. About 70%\n> > > efficient. During the compress phase all four cores were active until\n> > > the last 90 seconds. Writing the objects took over 23 minutes CPU\n> > > bound on one core.\n> > >\n> > > New pack file is: 270,594,853\n> > > Old one was: 344,543,752\n> > > It still has 828,660 objects\n> >\n> > You mean the pack for the gcc repo is now less than 300MB?  Wow.\n> >\n> >\n> > Nicolas\n> >\n>\n>\n> --\n> Jon Smirl\n> jonsmirl@gmail.com\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62656","messageId":"BAYC1-PASMTP08CFB6F824B1282649E5EAAE640@CEZ.ICE","threadId":"11182","inReplyTo":"9e4733910712102125w56c70c0cxb8b00a060b62077@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2007-12-11T06:01:30Z","receivedAt":"2007-12-11T06:01:30Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Tue, 11 Dec 2007 00:25:55 -0500\n\"Jon Smirl\" <jonsmirl@gmail.com> wrote:\n\n> Something is hurting bad with threads. 170 CPU minutes with one\n> thread, versus 195 CPU minutes with four threads.\n> \n> Is there a different memory allocator that can be used when\n> multithreaded on gcc? This whole problem may be coming from the memory\n> allocation function. git is hardly interacting at all on the thread\n> level so it's likely a problem in the C run-time.\n\nYou might want to try Google's malloc, it's basically a drop in replacement\nwith some optional built-in performance monitoring capabilities.  It is said\nto be much faster and better at threading than glibc's:\n\n  http://code.google.com/p/google-perftools/wiki/GooglePerformanceTools\n  http://google-perftools.googlecode.com/svn/trunk/doc/tcmalloc.html\n\n\nYou can LD_PRELOAD it or link directly.\n\nCheers,\nSean\n"},{"id":"62660","messageId":"9e4733910712102220u47601845q60ccfd754e71936b@mail.gmail.com","threadId":"11182","inReplyTo":"BAYC1-PASMTP08CFB6F824B1282649E5EAAE640@CEZ.ICE","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T06:20:37Z","receivedAt":"2007-12-11T06:20:37Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/11/07, Sean <seanlkml@sympatico.ca> wrote:\n> On Tue, 11 Dec 2007 00:25:55 -0500\n> \"Jon Smirl\" <jonsmirl@gmail.com> wrote:\n>\n> > Something is hurting bad with threads. 170 CPU minutes with one\n> > thread, versus 195 CPU minutes with four threads.\n> >\n> > Is there a different memory allocator that can be used when\n> > multithreaded on gcc? This whole problem may be coming from the memory\n> > allocation function. git is hardly interacting at all on the thread\n> > level so it's likely a problem in the C run-time.\n>\n> You might want to try Google's malloc, it's basically a drop in replacement\n> with some optional built-in performance monitoring capabilities.  It is said\n> to be much faster and better at threading than glibc's:\n>\n>   http://code.google.com/p/google-perftools/wiki/GooglePerformanceTools\n>   http://google-perftools.googlecode.com/svn/trunk/doc/tcmalloc.html\n>\n>\n> You can LD_PRELOAD it or link directly.\n\nI'm 45 minutes into a run using it. It doesn't seem to be any faster\nbut it is reducing memory consumption significantly. The run should be\ndone in another 20 minutes or so.\n\n\n>\n> Cheers,\n> Sean\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62672","messageId":"9e4733910712102301p5e6c4165v6afb32d157478828@mail.gmail.com","threadId":"11182","inReplyTo":"9e4733910712102129v140c2affqf2e73e75855b61ea@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T07:01:11Z","receivedAt":"2007-12-11T07:01:11Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Switching to the Google perftools malloc\nhttp://goog-perftools.sourceforge.net/\n\n10%   30  828M\n20%   15  831M\n30%   10  834M\n40%   50  1014M\n50%   80  1086M\n60%   80  1500M\n70% 200  1.53G\n80% 200  1.85G\n90% 260  1.87G\n95% 520  1.97G\n100% 1335 2.24G\n\nGoogle allocator knocked 600MB off from memory use.\nMemory consumption did not fall during the write out phase like it did with gcc.\n\nSince all of this is with the same code except for changing the\nthreading split, those runs where memory consumption went to 4.5GB\nwith the gcc allocator must have triggered an extreme problem with\nfragmentation.\n\nTotal CPU time 196 CPU minutes vs 190 for gcc. Google's claims of\nbeing faster are not true.\n\nSo why does our threaded code take 20 CPU minutes longer (12%) to run\nthan the same code with a single thread? Clock time is obviously\nfaster. Are the threads working too close to each other in memory and\nbouncing cache lines between the cores? Q6600 is just two E6600s in\nthe same package, the caches are not shared.\n\nWhy does the threaded code need 2.24GB (google allocator, 2.85GB gcc)\nwith 4 threads? But only need 950MB with one thread? Where's the extra\ngigabyte going?\n\nIs there another allocator to try? One that combines Google's\nefficiency with gcc's speed?\n\n\nOn 12/11/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> I added the gcc people to the CC, it's their repository. Maybe they\n> can help up sort this out.\n>\n> On 12/11/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > On 12/10/07, Nicolas Pitre <nico@cam.org> wrote:\n> > > On Mon, 10 Dec 2007, Jon Smirl wrote:\n> > >\n> > > > New run using same configuration. With the addition of the more\n> > > > efficient load balancing patches and delta cache accounting.\n> > > >\n> > > > Seconds are wall clock time. They are lower since the patch made\n> > > > threading better at using all four cores. I am stuck at 380-390% CPU\n> > > > utilization for the git process.\n> > > >\n> > > > complete seconds RAM\n> > > > 10%   60    900M (includes counting)\n> > > > 20%   15    900M\n> > > > 30%   15    900M\n> > > > 40%   50    1.2G\n> > > > 50%   80    1.3G\n> > > > 60%   70    1.7G\n> > > > 70%   140  1.8G\n> > > > 80%   180  2.0G\n> > > > 90%   280  2.2G\n> > > > 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> > > > 100% 1390 2.85G\n> > > > During the writing phase RAM fell to 1.6G\n> > > > What is being freed in the writing phase??\n> > >\n> > > The cached delta results, but you put a cap of 256MB for them.\n> > >\n> > > Could you try again with that cache disabled entirely, with\n> > > pack.deltacachesize = 1 (don't use 0 as that means unbounded).\n> > >\n> > > And then, while still keeping the delta cache disabled, could you try\n> > > with pack.threads = 2, and pack.threads = 1 ?\n> > >\n> > > I'm sorry to ask you to do this but I don't have enough ram to even\n> > > complete a repack with threads=2 so I'm reattempting single threaded at\n> > > the moment.  But I really wonder if the threading has such an effect on\n> > > memory usage.\n> >\n> > I already have a threads = 1 running with this config. Binary and\n> > config were same from threads=4 run.\n> >\n> > 10% 28min 950M\n> > 40% 135min 950M\n> > 50% 157min 900M\n> > 60% 160min 830M\n> > 100% 170min 830M\n> >\n> > Something is hurting bad with threads. 170 CPU minutes with one\n> > thread, versus 195 CPU minutes with four threads.\n> >\n> > Is there a different memory allocator that can be used when\n> > multithreaded on gcc? This whole problem may be coming from the memory\n> > allocation function. git is hardly interacting at all on the thread\n> > level so it's likely a problem in the C run-time.\n> >\n> > [core]\n> >         repositoryformatversion = 0\n> >         filemode = true\n> >         bare = false\n> >         logallrefupdates = true\n> > [pack]\n> >         threads = 1\n> >         deltacachesize = 256M\n> >         windowmemory = 256M\n> >         deltacachelimit = 0\n> > [remote \"origin\"]\n> >         url = git://git.infradead.org/gcc.git\n> >         fetch = +refs/heads/*:refs/remotes/origin/*\n> > [branch \"trunk\"]\n> >         remote = origin\n> >         merge = refs/heads/trunk\n> >\n> >\n> >\n> >\n> > >\n> > >\n> > >\n> > > >\n> > > > I have no explanation for the change in RAM usage. Two guesses come to\n> > > > mind. Memory fragmentation. Or the change in the way the work was\n> > > > split up altered RAM usage.\n> > > >\n> > > > Total CPU time was 195 minutes in 70 minutes clock time. About 70%\n> > > > efficient. During the compress phase all four cores were active until\n> > > > the last 90 seconds. Writing the objects took over 23 minutes CPU\n> > > > bound on one core.\n> > > >\n> > > > New pack file is: 270,594,853\n> > > > Old one was: 344,543,752\n> > > > It still has 828,660 objects\n> > >\n> > > You mean the pack for the gcc repo is now less than 300MB?  Wow.\n> > >\n> > >\n> > > Nicolas\n> > >\n> >\n> >\n> > --\n> > Jon Smirl\n> > jonsmirl@gmail.com\n> >\n>\n>\n> --\n> Jon Smirl\n> jonsmirl@gmail.com\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62679","messageId":"475E3D86.9030408@op5.se","threadId":"11182","inReplyTo":"9e4733910712102301p5e6c4165v6afb32d157478828@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-12-11T07:34:30Z","receivedAt":"2007-12-11T07:34:30Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Jon Smirl wrote:\n> Switching to the Google perftools malloc\n> http://goog-perftools.sourceforge.net/\n> \n> Google allocator knocked 600MB off from memory use.\n> Memory consumption did not fall during the write out phase like it did with gcc.\n> \n> Since all of this is with the same code except for changing the\n> threading split, those runs where memory consumption went to 4.5GB\n> with the gcc allocator must have triggered an extreme problem with\n> fragmentation.\n> \n> Total CPU time 196 CPU minutes vs 190 for gcc. Google's claims of\n> being faster are not true.\n> \n\nDid you use the tcmalloc with heap checker/profiler, or tcmalloc_minimal?\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"62699","messageId":"85bq8xlc8w.fsf@lola.goethe.zz","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712102225240.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-12-11T11:08:47Z","receivedAt":"2007-12-11T11:08:47Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Mon, 10 Dec 2007, Junio C Hamano wrote:\n>\n>> \"Jon Smirl\" <jonsmirl@gmail.com> writes:\n>> \n>> > 95%   530  2.8G - 1,420 total to here, previous was 1,983\n>> > 100% 1390 2.85G\n>> > During the writing phase RAM fell to 1.6G\n>> > What is being freed in the writing phase??\n>> \n>> entry->delta_data is the only thing I can think of that are freed\n>> in the function that have been allocated much earlier before entering\n>> the function.\n>\n> Yet all ->delta-data instances are limited to 256MB according to Jon's \n> config.\n\nMaybe address space fragmentation is involved here?  malloc/free for\nlarge areas works using mmap in glibc.  There must be enough\n_contiguous_ space for a new allocation to succeed.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"62703","messageId":"20071211120851.GK30948@artemis.madism.org","threadId":"11182","inReplyTo":"85bq8xlc8w.fsf@lola.goethe.zz","subject":"Re: Something is broken in repack","fromName":"Pierre Habouzit","fromEmail":"madcoder@artemis.madism.org","sentAt":"2007-12-11T12:08:51Z","receivedAt":"2007-12-11T12:08:51Z","isPatch":false,"sender":{"key":"madcoder@artemis.madism.org","avatar":null},"body":"On Tue, Dec 11, 2007 at 11:08:47AM +0000, David Kastrup wrote:\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > On Mon, 10 Dec 2007, Junio C Hamano wrote:\n> >\n> >> \"Jon Smirl\" <jonsmirl@gmail.com> writes:\n> >> \n> >> > 95%   530  2.8G - 1,420 total to here, previous was 1,983\n> >> > 100% 1390 2.85G\n> >> > During the writing phase RAM fell to 1.6G\n> >> > What is being freed in the writing phase??\n> >> \n> >> entry->delta_data is the only thing I can think of that are freed\n> >> in the function that have been allocated much earlier before entering\n> >> the function.\n> >\n> > Yet all ->delta-data instances are limited to 256MB according to Jon's \n> > config.\n> \n> Maybe address space fragmentation is involved here?  malloc/free for\n> large areas works using mmap in glibc.  There must be enough\n> _contiguous_ space for a new allocation to succeed.\n\n  Well, that's interesting, but there is a way to know for sure instead\nof taking bets. Just use valgrind --tool=massif and look at the pretty\npicture, it'll tell what was going on very accurately.\n\n  Note that I find your explanation unlikely: glibc uses mmap for sizes\nover 128k by default (IIRC), and as soon as you use mmaps, that's the\nkernel that deals with the address space, and it's not necessarily\ncontiguous, that's only true for the heap.\n-- \n·O·  Pierre Habouzit\n··O                                                madcoder@debian.org\nOOO                                                http://www.madism.org\n"},{"id":"62704","messageId":"851w9tl90i.fsf@lola.goethe.zz","threadId":"11182","inReplyTo":"20071211120851.GK30948@artemis.madism.org","subject":"Re: Something is broken in repack","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-12-11T12:18:37Z","receivedAt":"2007-12-11T12:18:37Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Pierre Habouzit <madcoder@artemis.madism.org> writes:\n\n> On Tue, Dec 11, 2007 at 11:08:47AM +0000, David Kastrup wrote:\n>\n>> Maybe address space fragmentation is involved here?  malloc/free for\n>> large areas works using mmap in glibc.  There must be enough\n>> _contiguous_ space for a new allocation to succeed.\n>\n>   Note that I find your explanation unlikely: glibc uses mmap for\n> sizes over 128k by default (IIRC), and as soon as you use mmaps,\n> that's the kernel that deals with the address space, and it's not\n> necessarily contiguous, that's only true for the heap.\n\nEvery single allocation needs to be contiguous in virtual address space\nand must not collide with existing virtual address space allocations.\nSo fragmentation is at least a logistical issue.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"62713","messageId":"alpine.LFD.0.99999.0712110828250.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712102129v140c2affqf2e73e75855b61ea@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T13:31:32Z","receivedAt":"2007-12-11T13:31:32Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Jon Smirl wrote:\n\n> I added the gcc people to the CC, it's their repository. Maybe they\n> can help up sort this out.\n\nUnless there is a Git expert amongst the gcc crowd, I somehow doubt it. \nAnd gcc people with an interest in Git internals are probably already on \nthe Git mailing list.\n\n\nNicolas\n"},{"id":"62716","messageId":"alpine.LFD.0.99999.0712110832251.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712102301p5e6c4165v6afb32d157478828@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T13:49:32Z","receivedAt":"2007-12-11T13:49:32Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Jon Smirl wrote:\n\n> Switching to the Google perftools malloc\n> http://goog-perftools.sourceforge.net/\n> \n> 10%   30  828M\n> 20%   15  831M\n> 30%   10  834M\n> 40%   50  1014M\n> 50%   80  1086M\n> 60%   80  1500M\n> 70% 200  1.53G\n> 80% 200  1.85G\n> 90% 260  1.87G\n> 95% 520  1.97G\n> 100% 1335 2.24G\n> \n> Google allocator knocked 600MB off from memory use.\n> Memory consumption did not fall during the write out phase like it did with gcc.\n> \n> Since all of this is with the same code except for changing the\n> threading split, those runs where memory consumption went to 4.5GB\n> with the gcc allocator must have triggered an extreme problem with\n> fragmentation.\n\nDid you mean the glibc allocator?\n\n> Total CPU time 196 CPU minutes vs 190 for gcc. Google's claims of\n> being faster are not true.\n> \n> So why does our threaded code take 20 CPU minutes longer (12%) to run\n> than the same code with a single thread? Clock time is obviously\n> faster. Are the threads working too close to each other in memory and\n> bouncing cache lines between the cores? Q6600 is just two E6600s in\n> the same package, the caches are not shared.\n\nOf course there'll always be a certain amount of wasted cycles when \nthreaded.  The locking overhead, the extra contention for IO, etc.  So \n12% overhead (3% per thread) when using 4 threads is not that bad I \nwould say.\n\n> Why does the threaded code need 2.24GB (google allocator, 2.85GB gcc)\n> with 4 threads? But only need 950MB with one thread? Where's the extra\n> gigabyte going?\n\nI really don't know.\n\nDid you try with pack.deltacachesize set to 1 ?\n\nAnd yet, this is still missing the actual issue.  The issue being that \nthe 2.1GB pack as a _source_ doesn't cause as much memory to be \nallocated even if the _result_ pack ends up being the same.\n\nI was able to repack the 2.1GB pack on my machine which has 1GB of ram. \nNow that it has been repacked, I can't repack it anymore, even when \nsingle threaded, as it start crowling into swap fairly quickly.  It is \nreally non intuitive and actually senseless that Git would require twice \nas much RAM to deal with a pack that is 7 times smaller.\n\n\nNicolas (still puzzled)\n"},{"id":"62725","messageId":"alpine.LFD.0.99999.0712110951070.555@xanadu.home","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712110832251.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T15:00:56Z","receivedAt":"2007-12-11T15:00:56Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Nicolas Pitre wrote:\n\n> And yet, this is still missing the actual issue.  The issue being that \n> the 2.1GB pack as a _source_ doesn't cause as much memory to be \n> allocated even if the _result_ pack ends up being the same.\n> \n> I was able to repack the 2.1GB pack on my machine which has 1GB of ram. \n> Now that it has been repacked, I can't repack it anymore, even when \n> single threaded, as it start crowling into swap fairly quickly.  It is \n> really non intuitive and actually senseless that Git would require twice \n> as much RAM to deal with a pack that is 7 times smaller.\n\nOK, here's something else for you to try:\n\n\tcore.deltabasecachelimit=0\n\tpack.threads=2\n\tpack.deltacachesize=1\n\nWith that I'm able to repack the small gcc pack on my machine with 1GB \nof ram using:\n\n\tgit repack -a -f -d --window=250 --depth=250\n\nand top reports a ~700m virt and ~500m res without hitting swap at all.\nIt is only at 25% so far, but I was unable to get that far before.\n\nWould be curious to know what you get with 4 threads on your machine.\n\n\nNicolas\n"},{"id":"62735","messageId":"9e4733910712110736w34495ba2l86b2de82055620fd@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712110951070.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T15:36:20Z","receivedAt":"2007-12-11T15:36:20Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/11/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 11 Dec 2007, Nicolas Pitre wrote:\n>\n> > And yet, this is still missing the actual issue.  The issue being that\n> > the 2.1GB pack as a _source_ doesn't cause as much memory to be\n> > allocated even if the _result_ pack ends up being the same.\n> >\n> > I was able to repack the 2.1GB pack on my machine which has 1GB of ram.\n> > Now that it has been repacked, I can't repack it anymore, even when\n> > single threaded, as it start crowling into swap fairly quickly.  It is\n> > really non intuitive and actually senseless that Git would require twice\n> > as much RAM to deal with a pack that is 7 times smaller.\n>\n> OK, here's something else for you to try:\n>\n>         core.deltabasecachelimit=0\n>         pack.threads=2\n>         pack.deltacachesize=1\n>\n> With that I'm able to repack the small gcc pack on my machine with 1GB\n> of ram using:\n>\n>         git repack -a -f -d --window=250 --depth=250\n>\n> and top reports a ~700m virt and ~500m res without hitting swap at all.\n> It is only at 25% so far, but I was unable to get that far before.\n>\n> Would be curious to know what you get with 4 threads on your machine.\n\nChanging those parameters really slowed down counting the objects. I\nused to be able to count in 45 seconds now it took 130 seconds. I am\nstill have the Google allocator linked in.\n\n4 threads, cumulative clock time\n25%     200 seconds, 820/627M\n55%     510 seconds, 1240/1000M - little late recording\n75%     15 minutes, 1658/1500M\n90%      22 minutes, 1974/1800M\nit's still running but there is no significant change.\n\nAre two types of allocations being mixed?\n1) long term, global objects kept until the end of everything\n2) volatile, private objects allocated only while the object is being\ncompressed and then freed\n\nSeparating these would make a big difference to the fragmentation\nproblem. Single threading probably wouldn't see a fragmentation\nproblem from mixing the allocation types.\n\nWhen a thread is created it could allocated a private 20MB (or\nwhatever) pool. The volatile, private objects would come from that\npool. Long term objects would stay in the global pool. Since they are\nlong term they will just get laid down sequentially in memory.\nSeparating these allocation types make things way easier for malloc.\n\nCPU time would be helped by removing some of the locking if possible.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62741","messageId":"alpine.LFD.0.99999.0712111117440.555@xanadu.home","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712110951070.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T16:20:17Z","receivedAt":"2007-12-11T16:20:17Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Nicolas Pitre wrote:\n\n> OK, here's something else for you to try:\n> \n> \tcore.deltabasecachelimit=0\n> \tpack.threads=2\n> \tpack.deltacachesize=1\n> \n> With that I'm able to repack the small gcc pack on my machine with 1GB \n> of ram using:\n> \n> \tgit repack -a -f -d --window=250 --depth=250\n> \n> and top reports a ~700m virt and ~500m res without hitting swap at all.\n> It is only at 25% so far, but I was unable to get that far before.\n\nWell, around 55% memory usage skyrocketed to 1.6GB and the system went \ndeep into swap.  So I restarted it with no threads.\n\nNicolas (even more puzzled)\n"},{"id":"62742","messageId":"9e4733910712110821o7748802ag75d9df4be8b2c123@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712111117440.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T16:21:57Z","receivedAt":"2007-12-11T16:21:57Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/11/07, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 11 Dec 2007, Nicolas Pitre wrote:\n>\n> > OK, here's something else for you to try:\n> >\n> >       core.deltabasecachelimit=0\n> >       pack.threads=2\n> >       pack.deltacachesize=1\n> >\n> > With that I'm able to repack the small gcc pack on my machine with 1GB\n> > of ram using:\n> >\n> >       git repack -a -f -d --window=250 --depth=250\n> >\n> > and top reports a ~700m virt and ~500m res without hitting swap at all.\n> > It is only at 25% so far, but I was unable to get that far before.\n>\n> Well, around 55% memory usage skyrocketed to 1.6GB and the system went\n> deep into swap.  So I restarted it with no threads.\n>\n> Nicolas (even more puzzled)\n\nOn the plus side you are seeing what I see, so it proves I am not imagining it.\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62744","messageId":"alpine.LFD.0.9999.0712110806540.25032@woody.linux-foundation.org","threadId":"11182","inReplyTo":"9e4733910712102301p5e6c4165v6afb32d157478828@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T16:33:21Z","receivedAt":"2007-12-11T16:33:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Jon Smirl wrote:\n> \n> So why does our threaded code take 20 CPU minutes longer (12%) to run\n> than the same code with a single thread?\n\nThreaded code *always* takes more CPU time. The only thing you can hope \nfor is a wall-clock reduction. You're seeing probably a combination of \n (a) more cache misses\n (b) bigger dataset active at a time\nand a probably fairly miniscule\n (c) threading itself tends to have some overheads.\n\n> Q6600 is just two E6600s in the same package, the caches are not shared.\n\nSure they are shared. They're just not *entirely* shared. But they are \nshared between each two cores, so each thread essentially has only half \nthe cache they had with the non-threaded version.\n\nThreading is *not* a magic solution to all problems. It gives you \npotentially twice the CPU power, but there are real downsides that you \nshould keep in mind.\n\n> Why does the threaded code need 2.24GB (google allocator, 2.85GB gcc)\n> with 4 threads? But only need 950MB with one thread? Where's the extra\n> gigabyte going?\n\nI suspect that it's really simple: you have a few rather big files in the \ngcc history, with deep delta chains. And what happens when you have four \nthreads running at the same time is that they all need to keep all those \nobjects that they are working on - and their hash state - in memory at the \nsame time!\n\nSo if you want to use more threads, that _forces_ you to have a bigger \nmemory footprint, simply because you have more \"live\" objects that you \nwork on. Normally, that isn't much of a problem, since most source files \nare small, but if you have a few deep delta chains on big files, both the \ndelta chain itself is going to use memory (you may have limited the size \nof the cache, but it's still needed for the actual delta generation, so \nit's not like the memory usage went away).\n\nThat said, I suspect there are a few things fighting you:\n\n - threading is hard. I haven't looked a lot at the changes Nico did to do \n   a threaded object packer, but what I've seen does not convince me it is \n   correct. The \"trg_entry\" accesses are *mostly* protected with \n   \"cache_lock\", but nothing else really seems to be, so quite frankly, I \n   wouldn't trust the threaded version very much. It's off by default, and \n   for a good reason, I think.\n\n   For example: the packing code does this:\n\n\tif (!src->data) {\n\t\tread_lock();\n\t\tsrc->data = read_sha1_file(src_entry->idx.sha1, &type, &sz);\n\t\tread_unlock();\n\t\t...\n\n   and that's racy. If two threads come in at roughly the same time and \n   see a NULL src->data, theÿ́'ll both get the lock, and they'll both \n   (serially) try to fill it in. It will all *work*, but one of them will \n   have done unnecessary work, and one of them will have their result \n   thrown away and leaked.\n\n   Are you hitting issues like this? I dunno. The object sorting means \n   that different threads normally shouldn't look at the same objects (not \n   even the sources), so probably not, but basically, I wouldn't trust the \n   threading 100%. It needs work, and it needs to stay off by default.\n\n - you're working on a problem that isn't really even worth optimizing \n   that much. The *normal* case is to re-use old deltas, which makes all \n   of the issues you are fighting basically go away (because you only have \n   a few _incremental_ objects that need deltaing). \n\n   In other words: the _real_ optimizations have already been done, and \n   are done elsewhere, and are much smarter (the best way to optimize X is \n   not to make X run fast, but to avoid doing X in the first place!). The \n   thing you are trying to work with is the one-time-only case where you \n   explicitly disable that big and important optimization, and then you \n   complain about the end result being slow!\n\n   It's like saying that you're compiling with extreme debugging and no\n   optimizations, and then complaining that the end result doesn't run as \n   fast as if you used -O2. Except this is a hundred times worse, because \n   you literally asked git to do the really expensive thing that it really \n   really doesn't want to do ;)\n\n> Is there another allocator to try? One that combines Google's\n> efficiency with gcc's speed?\n\nSee above: I'd look around at threading-related bugs and check the way we \nlock (or don't) accesses.\n\n\t\tLinus\n"},{"id":"62747","messageId":"alpine.LFD.0.99999.0712111202470.555@xanadu.home","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712110806540.25032@woody.linux-foundation.org","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T17:21:11Z","receivedAt":"2007-12-11T17:21:11Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Linus Torvalds wrote:\n\n> That said, I suspect there are a few things fighting you:\n> \n>  - threading is hard. I haven't looked a lot at the changes Nico did to do \n>    a threaded object packer, but what I've seen does not convince me it is \n>    correct. The \"trg_entry\" accesses are *mostly* protected with \n>    \"cache_lock\", but nothing else really seems to be, so quite frankly, I \n>    wouldn't trust the threaded version very much. It's off by default, and \n>    for a good reason, I think.\n\nI beg to differ (of course, since I always know precisely what I do, and \nlike you, my code never has bugs).\n\nSeriously though, the trg_entry has not to be protected at all.  Why? \nSimply because each thread has its own exclusive set of objects which no \nother threads ever mess with.  They never overlap.\n\n>    For example: the packing code does this:\n> \n> \tif (!src->data) {\n> \t\tread_lock();\n> \t\tsrc->data = read_sha1_file(src_entry->idx.sha1, &type, &sz);\n> \t\tread_unlock();\n> \t\t...\n> \n>    and that's racy. If two threads come in at roughly the same time and \n>    see a NULL src->data, theÿ́'ll both get the lock, and they'll both \n>    (serially) try to fill it in. It will all *work*, but one of them will \n>    have done unnecessary work, and one of them will have their result \n>    thrown away and leaked.\n\nNo.  Once again, it is impossible for two threads to ever see the same \nsrc->data at all.  The lock is there simply because read_sha1_file() is \nnot reentrant.\n\n>    Are you hitting issues like this? I dunno. The object sorting means \n>    that different threads normally shouldn't look at the same objects (not \n>    even the sources), so probably not, but basically, I wouldn't trust the \n>    threading 100%. It needs work, and it needs to stay off by default.\n\nFor now it is, but I wouldn't say it really needs significant work at \nthis point.  The latest thread patches were more about tuning than \ncorrectness.\n\nWhat the threading could be doing, though, is uncovering some other \nbugs, like in the pack mmap windowing code for example.  Although that \ncode is serialized by the read lock above, the fact that multiple \nthreads are hammering on it in turns means that the mmap window is \npossibly seeking back and forth much more often than otherwise, possibly \nleaking something in the process.\n\n>  - you're working on a problem that isn't really even worth optimizing \n>    that much. The *normal* case is to re-use old deltas, which makes all \n>    of the issues you are fighting basically go away (because you only have \n>    a few _incremental_ objects that need deltaing). \n> \n>    In other words: the _real_ optimizations have already been done, and \n>    are done elsewhere, and are much smarter (the best way to optimize X is \n>    not to make X run fast, but to avoid doing X in the first place!). The \n>    thing you are trying to work with is the one-time-only case where you \n>    explicitly disable that big and important optimization, and then you \n>    complain about the end result being slow!\n> \n>    It's like saying that you're compiling with extreme debugging and no\n>    optimizations, and then complaining that the end result doesn't run as \n>    fast as if you used -O2. Except this is a hundred times worse, because \n>    you literally asked git to do the really expensive thing that it really \n>    really doesn't want to do ;)\n\nLinus, please pay attention to the _actual_ important issue here.\n\nSure I've been tuning the threading code in parallel to the attempt to \ndebug this memory usage issue.\n\nBUT.  The point is that repacking the gcc repo using \"git repack -a -f \n--window=250\" has a radically different memory usage profile whether you \ndo the repack on the earlier 2.1GB pack or the later 300MB pack.  \n_That_ is the issue.  Ironically, it is the 300MB pack that causes the \nrepack to blow memory usage out of proportion.\n\nAnd in both cases, the threading code has to do the same \nwork whether or not the original pack was densely packed or not since -f \nthrows away every existing deltas anyway.\n\nSo something is fishy elsewhere than in the packing code.\n\n\nNicolas\n"},{"id":"62748","messageId":"20071211.092402.266823343.davem@davemloft.net","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712111202470.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-11T17:24:02Z","receivedAt":"2007-12-11T17:24:02Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: Nicolas Pitre <nico@cam.org>\nDate: Tue, 11 Dec 2007 12:21:11 -0500 (EST)\n\n> BUT.  The point is that repacking the gcc repo using \"git repack -a -f \n> --window=250\" has a radically different memory usage profile whether you \n> do the repack on the earlier 2.1GB pack or the later 300MB pack.  \n\nIf you repack on the smaller pack file, git has to expand more stuff\ninternally in order to search the deltas, whereas with the larger pack\nfile I bet git has to less often undelta'ify to get base objects blobs\nfor delta search.\n\nIn fact that behavior makes perfect sense to me and I don't understand\nGIT internals very well :-)\n"},{"id":"62749","messageId":"4aca3dc20712110928ybb84c16n40b6dbd50feddb06@mail.gmail.com","threadId":"11182","inReplyTo":"9e4733910712102301p5e6c4165v6afb32d157478828@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Daniel Berlin","fromEmail":"dberlin@dberlin.org","sentAt":"2007-12-11T17:28:25Z","receivedAt":"2007-12-11T17:28:25Z","isPatch":false,"sender":{"key":"dberlin@dberlin.org","avatar":null},"body":"On 12/11/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n>\n> Total CPU time 196 CPU minutes vs 190 for gcc. Google's claims of\n> being faster are not true.\n\nDepends on your allocation patterns. For our apps, it certainly is :)\nOf course, i don't know if we've updated the external allocator in a\nwhile, i'll bug the people in charge of it.\n"},{"id":"62752","messageId":"alpine.LFD.0.99999.0712111237530.555@xanadu.home","threadId":"11182","inReplyTo":"20071211.092402.266823343.davem@davemloft.net","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T17:44:47Z","receivedAt":"2007-12-11T17:44:47Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, David Miller wrote:\n\n> From: Nicolas Pitre <nico@cam.org>\n> Date: Tue, 11 Dec 2007 12:21:11 -0500 (EST)\n> \n> > BUT.  The point is that repacking the gcc repo using \"git repack -a -f \n> > --window=250\" has a radically different memory usage profile whether you \n> > do the repack on the earlier 2.1GB pack or the later 300MB pack.  \n> \n> If you repack on the smaller pack file, git has to expand more stuff\n> internally in order to search the deltas, whereas with the larger pack\n> file I bet git has to less often undelta'ify to get base objects blobs\n> for delta search.\n\nOf course.  I came to that conclusion two days ago.  And despite being \npretty familiar with the involved code (I wrote part of it myself) I \njust can't spot anything wrong with it so far.\n\nBut somehow the threading code keep distracting people from that issue \nsince it gets to do the same work whether or not the source pack is \ndensely packed or not.\n\nNicolas \n(who wish he had access to a much faster machine to investigate this issue)\n"},{"id":"62763","messageId":"9e4733910712111043h6a361996x740f4dba3d742da5@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712110806540.25032@woody.linux-foundation.org","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-11T18:43:19Z","receivedAt":"2007-12-11T18:43:19Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/11/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Tue, 11 Dec 2007, Jon Smirl wrote:\n> >\n> > So why does our threaded code take 20 CPU minutes longer (12%) to run\n> > than the same code with a single thread?\n>\n> Threaded code *always* takes more CPU time. The only thing you can hope\n> for is a wall-clock reduction. You're seeing probably a combination of\n>  (a) more cache misses\n>  (b) bigger dataset active at a time\n> and a probably fairly miniscule\n>  (c) threading itself tends to have some overheads.\n>\n> > Q6600 is just two E6600s in the same package, the caches are not shared.\n>\n> Sure they are shared. They're just not *entirely* shared. But they are\n> shared between each two cores, so each thread essentially has only half\n> the cache they had with the non-threaded version.\n>\n> Threading is *not* a magic solution to all problems. It gives you\n> potentially twice the CPU power, but there are real downsides that you\n> should keep in mind.\n>\n> > Why does the threaded code need 2.24GB (google allocator, 2.85GB gcc)\n> > with 4 threads? But only need 950MB with one thread? Where's the extra\n> > gigabyte going?\n>\n> I suspect that it's really simple: you have a few rather big files in the\n> gcc history, with deep delta chains. And what happens when you have four\n> threads running at the same time is that they all need to keep all those\n> objects that they are working on - and their hash state - in memory at the\n> same time!\n>\n> So if you want to use more threads, that _forces_ you to have a bigger\n> memory footprint, simply because you have more \"live\" objects that you\n> work on. Normally, that isn't much of a problem, since most source files\n> are small, but if you have a few deep delta chains on big files, both the\n> delta chain itself is going to use memory (you may have limited the size\n> of the cache, but it's still needed for the actual delta generation, so\n> it's not like the memory usage went away).\n\nThis makes sense. Those runs that blew up to 4.5GB were a combination\nof this effect and fragmentation in the gcc allocator. Google\nallocator appears to be much better at controlling fragmentation.\n\nIs there a reasonable scheme to force the chains to only be loaded\nonce and then shared between worker threads? The memory blow up\nappears to be directly correlated with chain length.\n\n>\n> That said, I suspect there are a few things fighting you:\n>\n>  - threading is hard. I haven't looked a lot at the changes Nico did to do\n>    a threaded object packer, but what I've seen does not convince me it is\n>    correct. The \"trg_entry\" accesses are *mostly* protected with\n>    \"cache_lock\", but nothing else really seems to be, so quite frankly, I\n>    wouldn't trust the threaded version very much. It's off by default, and\n>    for a good reason, I think.\n>\n>    For example: the packing code does this:\n>\n>         if (!src->data) {\n>                 read_lock();\n>                 src->data = read_sha1_file(src_entry->idx.sha1, &type, &sz);\n>                 read_unlock();\n>                 ...\n>\n>    and that's racy. If two threads come in at roughly the same time and\n>    see a NULL src->data, theÿ́'ll both get the lock, and they'll both\n>    (serially) try to fill it in. It will all *work*, but one of them will\n>    have done unnecessary work, and one of them will have their result\n>    thrown away and leaked.\n\nThat may account for the threaded version needing an extra 20 minutes\nCPU time.  An extra 12% of CPU seems like too much overhead for\nthreading. Just letting a couple of those long chain compressions be\ndone twice\n\n>\n>    Are you hitting issues like this? I dunno. The object sorting means\n>    that different threads normally shouldn't look at the same objects (not\n>    even the sources), so probably not, but basically, I wouldn't trust the\n>    threading 100%. It needs work, and it needs to stay off by default.\n>\n>  - you're working on a problem that isn't really even worth optimizing\n>    that much. The *normal* case is to re-use old deltas, which makes all\n>    of the issues you are fighting basically go away (because you only have\n>    a few _incremental_ objects that need deltaing).\n\nI agree, this problem only occurs when people import giant\nrepositories. But every time someone hits these problems they declare\ngit to be screwed up and proceed to thrash it in their blogs.\n\n>    In other words: the _real_ optimizations have already been done, and\n>    are done elsewhere, and are much smarter (the best way to optimize X is\n>    not to make X run fast, but to avoid doing X in the first place!). The\n>    thing you are trying to work with is the one-time-only case where you\n>    explicitly disable that big and important optimization, and then you\n>    complain about the end result being slow!\n>\n>    It's like saying that you're compiling with extreme debugging and no\n>    optimizations, and then complaining that the end result doesn't run as\n>    fast as if you used -O2. Except this is a hundred times worse, because\n>    you literally asked git to do the really expensive thing that it really\n>    really doesn't want to do ;)\n>\n> > Is there another allocator to try? One that combines Google's\n> > efficiency with gcc's speed?\n>\n> See above: I'd look around at threading-related bugs and check the way we\n> lock (or don't) accesses.\n>\n>                 Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62764","messageId":"alpine.LFD.0.99999.0712111349270.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712111043h6a361996x740f4dba3d742da5@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-11T18:57:22Z","receivedAt":"2007-12-11T18:57:22Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Jon Smirl wrote:\n\n> This makes sense. Those runs that blew up to 4.5GB were a combination\n> of this effect and fragmentation in the gcc allocator.\n\nI disagree.  This is insane.\n\n> Google allocator appears to be much better at controlling fragmentation.\n\nIndeed.  And if fragmentation is indeed wasting half of Git's memory \nusage then we'll have to come with a custom memory allocator.\n\n> Is there a reasonable scheme to force the chains to only be loaded\n> once and then shared between worker threads? The memory blow up\n> appears to be directly correlated with chain length.\n\nNo.  That would be the equivalent of holding each revision of all files \nuncompressed all at once in memory.\n\n> > That said, I suspect there are a few things fighting you:\n> >\n> >  - threading is hard. I haven't looked a lot at the changes Nico did to do\n> >    a threaded object packer, but what I've seen does not convince me it is\n> >    correct. The \"trg_entry\" accesses are *mostly* protected with\n> >    \"cache_lock\", but nothing else really seems to be, so quite frankly, I\n> >    wouldn't trust the threaded version very much. It's off by default, and\n> >    for a good reason, I think.\n> >\n> >    For example: the packing code does this:\n> >\n> >         if (!src->data) {\n> >                 read_lock();\n> >                 src->data = read_sha1_file(src_entry->idx.sha1, &type, &sz);\n> >                 read_unlock();\n> >                 ...\n> >\n> >    and that's racy. If two threads come in at roughly the same time and\n> >    see a NULL src->data, theÿ́'ll both get the lock, and they'll both\n> >    (serially) try to fill it in. It will all *work*, but one of them will\n> >    have done unnecessary work, and one of them will have their result\n> >    thrown away and leaked.\n> \n> That may account for the threaded version needing an extra 20 minutes\n> CPU time.  An extra 12% of CPU seems like too much overhead for\n> threading. Just letting a couple of those long chain compressions be\n> done twice\n\nNo it may not.  This theory is wrong as explained before.\n\n> >\n> >    Are you hitting issues like this? I dunno. The object sorting means\n> >    that different threads normally shouldn't look at the same objects (not\n> >    even the sources), so probably not, but basically, I wouldn't trust the\n> >    threading 100%. It needs work, and it needs to stay off by default.\n> >\n> >  - you're working on a problem that isn't really even worth optimizing\n> >    that much. The *normal* case is to re-use old deltas, which makes all\n> >    of the issues you are fighting basically go away (because you only have\n> >    a few _incremental_ objects that need deltaing).\n> \n> I agree, this problem only occurs when people import giant\n> repositories. But every time someone hits these problems they declare\n> git to be screwed up and proceed to thrash it in their blogs.\n\nIt's not only for repack.  Someone just reported git-blame being \nunusable too due to insane memory usage, which I suspect is due to the \nsame issue.\n\n\nNicolas\n"},{"id":"62772","messageId":"alpine.LFD.0.9999.0712111055590.25032@woody.linux-foundation.org","threadId":"11182","inReplyTo":"9e4733910712111043h6a361996x740f4dba3d742da5@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-11T19:17:08Z","receivedAt":"2007-12-11T19:17:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 11 Dec 2007, Jon Smirl wrote:\n> >\n> > So if you want to use more threads, that _forces_ you to have a bigger\n> > memory footprint, simply because you have more \"live\" objects that you\n> > work on. Normally, that isn't much of a problem, since most source files\n> > are small, but if you have a few deep delta chains on big files, both the\n> > delta chain itself is going to use memory (you may have limited the size\n> > of the cache, but it's still needed for the actual delta generation, so\n> > it's not like the memory usage went away).\n> \n> This makes sense. Those runs that blew up to 4.5GB were a combination\n> of this effect and fragmentation in the gcc allocator. Google\n> allocator appears to be much better at controlling fragmentation.\n\nYes. I think we do have some case where we simply keep a lot of objects \naround, and if we are talking reasonably large deltas, we'll have the \nwhole delta-chain in memory just to unpack one single object.\n\nThe delta cache size limits kick in only when we explicitly cache old \ndelta results (in case they will be re-used, which is rather common), it \ndoesn't affect the normal \"I'm using this data right now\" case at all.\n\nAnd then fragmentation makes it much much worse. Since the allocation \npatterns aren't nice (they are pretty random and depend on just the sizes \nof the objects), and the lifetimes aren't always nicely nested _either_ \n(they become more so when you disable the cache entirely, but that's just \ndeath for performance), I'm not surprised that there can be memory \nallocators that end up having some issues.\n\n> Is there a reasonable scheme to force the chains to only be loaded\n> once and then shared between worker threads? The memory blow up\n> appears to be directly correlated with chain length.\n\nThe worker threads explicitly avoid touching the same objects, and no, you \ndefinitely don't want to explode the chains globally once, because the \nwhole point is that we do fit 15 years worth of history into 300MB of \npack-file thanks to having a very dense representation. The \"loaded once\" \npart is the mmap'ing of the pack-file into memory, but if you were to \nactually then try to expand the chains, you'd be talking about many *many* \nmore gigabytes of memory than you already see used ;)\n\nSo what you actually want to do is to just re-use already packed delta \nchains directly, which is what we normally do. But you are explicitly \nlooking at the \"--no-reuse-delta\" (aka \"git repack -f\") case, which is why \nit then blows up.\n\nI'm sure we can find places to improve. But I would like to re-iterate the \nstatement that you're kind of doing a \"don't do that then\" case which is \nreally - by design - meant to be done once and never again, and is using \nresources - again, pretty much by design - wildly inappropriately just to \nget an initial packing done.\n\n> That may account for the threaded version needing an extra 20 minutes\n> CPU time.  An extra 12% of CPU seems like too much overhead for\n> threading. Just letting a couple of those long chain compressions be\n> done twice\n\nWell, Nico pointed out that those things should all be thread-private \ndata, so no, the race isn't there (unless there's some other bug there).\n\n> I agree, this problem only occurs when people import giant\n> repositories. But every time someone hits these problems they declare\n> git to be screwed up and proceed to thrash it in their blogs.\n\nSure. I'd love to do global packing without paying the cost, but it really \nwas a design decision. Thanks to doing off-line packing (\"let it run \novernight on some beefy machine\") we can get better results. It's \nexpensive, yes. But it was pretty much meant to be expensive. It's a very \nefficient compression algorithm, after all, and you're turning it up to \neleven ;)\n\nI also suspect that the gcc archive makes things more interesting thanks \nto having some rather large files. The ChangeLog is probably the worst \ncase (large file with *lots* of edits), but I suspect the *.po files \naren't wonderful either.\n\n\t\t\tLinus\n"},{"id":"62781","messageId":"7v7ijldnq1.fsf@gitster.siamese.dyndns.org","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712111055590.25032@woody.linux-foundation.org","subject":"Re: Something is broken in repack","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-11T19:40:22Z","receivedAt":"2007-12-11T19:40:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Tue, 11 Dec 2007, Jon Smirl wrote:\n>> >\n>> > So if you want to use more threads, that _forces_ you to have a bigger\n>> > memory footprint, simply because you have more \"live\" objects that you\n>> > work on. Normally, that isn't much of a problem, since most source files\n>> > are small, but if you have a few deep delta chains on big files, both the\n>> > delta chain itself is going to use memory (you may have limited the size\n>> > of the cache, but it's still needed for the actual delta generation, so\n>> > it's not like the memory usage went away).\n>> \n>> This makes sense. Those runs that blew up to 4.5GB were a combination\n>> of this effect and fragmentation in the gcc allocator. Google\n>> allocator appears to be much better at controlling fragmentation.\n>\n> Yes. I think we do have some case where we simply keep a lot of objects \n> around, and if we are talking reasonably large deltas, we'll have the \n> whole delta-chain in memory just to unpack one single object.\n\nEh, excuse me.  unpack_delta_entry()\n\n - first unpacks the base object (this goes recursive);\n - uncompresses the delta;\n - applies the delta to the base to obtain the target object;\n - frees delta;\n - frees (but allows it to be cached) the base object;\n - returns the result\n\nSo no matter how deep a chain is, you keep only one delta at a time in\ncore, not whole delta-chain in core.\n\n> So what you actually want to do is to just re-use already packed delta \n> chains directly, which is what we normally do. But you are explicitly \n> looking at the \"--no-reuse-delta\" (aka \"git repack -f\") case, which is why \n> it then blows up.\n\nWhile that does not explain, as Nico pointed out, the huge difference\nbetween the two repack runs that have different starting pack, I would\nsay it is a fair thing to say.  If you have a suboptimal pack (i.e. not\nenough reusable deltas, as in the 2.1GB pack case), do run \"repack -f\",\nbut if you have a good pack (i.e. 300MB pack), don't.\n"},{"id":"62792","messageId":"475EF27B.7060609@op5.se","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712111237530.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-12-11T20:26:35Z","receivedAt":"2007-12-11T20:26:35Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n> On Tue, 11 Dec 2007, David Miller wrote:\n> \n>> From: Nicolas Pitre <nico@cam.org>\n>> Date: Tue, 11 Dec 2007 12:21:11 -0500 (EST)\n>>\n>>> BUT.  The point is that repacking the gcc repo using \"git repack -a -f \n>>> --window=250\" has a radically different memory usage profile whether you \n>>> do the repack on the earlier 2.1GB pack or the later 300MB pack.  \n>> If you repack on the smaller pack file, git has to expand more stuff\n>> internally in order to search the deltas, whereas with the larger pack\n>> file I bet git has to less often undelta'ify to get base objects blobs\n>> for delta search.\n> \n> Of course.  I came to that conclusion two days ago.  And despite being \n> pretty familiar with the involved code (I wrote part of it myself) I \n> just can't spot anything wrong with it so far.\n> \n> But somehow the threading code keep distracting people from that issue \n> since it gets to do the same work whether or not the source pack is \n> densely packed or not.\n> \n> Nicolas \n> (who wish he had access to a much faster machine to investigate this issue)\n\nIf it's still an issue next week, we'll have a 16 core (8 dual-core cpu's)\nmachine with some 32gb of ram in that'll be free for about two days.\nYou'll have to remind me about it though, as I've got a lot on my mind\nthese days.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"62796","messageId":"475EF453.90404@op5.se","threadId":"11182","inReplyTo":"7v7ijldnq1.fsf@gitster.siamese.dyndns.org","subject":"Re: Something is broken in repack","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-12-11T20:34:27Z","receivedAt":"2007-12-11T20:34:27Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n>> So what you actually want to do is to just re-use already packed delta \n>> chains directly, which is what we normally do. But you are explicitly \n>> looking at the \"--no-reuse-delta\" (aka \"git repack -f\") case, which is why \n>> it then blows up.\n> \n> While that does not explain, as Nico pointed out, the huge difference\n> between the two repack runs that have different starting pack, I would\n> say it is a fair thing to say.  If you have a suboptimal pack (i.e. not\n> enough reusable deltas, as in the 2.1GB pack case), do run \"repack -f\",\n> but if you have a good pack (i.e. 300MB pack), don't.\n\n\nI think this is too much of a mystery for a lot of people to let it go.\nEven I started looking into it, and I've got so little spare time just\nnow that I wouldn't stand much of a chance of making a contribution\neven if I had written the code originally.\n\nThat being said, I the fact that some git repositories really *can't*\nbe repacked on some machines (because it eats ALL virtual memory) is\nreally something that lowers git's reputation among huge projects.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"62853","messageId":"alpine.LFD.0.99999.0712112057390.555@xanadu.home","threadId":"11182","inReplyTo":"9e4733910712110821o7748802ag75d9df4be8b2c123@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-12T05:12:57Z","receivedAt":"2007-12-12T05:12:57Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 11 Dec 2007, Jon Smirl wrote:\n\n> On 12/11/07, Nicolas Pitre <nico@cam.org> wrote:\n> > On Tue, 11 Dec 2007, Nicolas Pitre wrote:\n> >\n> > > OK, here's something else for you to try:\n> > >\n> > >       core.deltabasecachelimit=0\n> > >       pack.threads=2\n> > >       pack.deltacachesize=1\n> > >\n> > > With that I'm able to repack the small gcc pack on my machine with 1GB\n> > > of ram using:\n> > >\n> > >       git repack -a -f -d --window=250 --depth=250\n> > >\n> > > and top reports a ~700m virt and ~500m res without hitting swap at all.\n> > > It is only at 25% so far, but I was unable to get that far before.\n> >\n> > Well, around 55% memory usage skyrocketed to 1.6GB and the system went\n> > deep into swap.  So I restarted it with no threads.\n> >\n> > Nicolas (even more puzzled)\n> \n> On the plus side you are seeing what I see, so it proves I am not imagining it.\n\nWell... This is weird.\n\nIt seems that memory fragmentation is really really killing us here.  \nThe fact that the Google allocator did manage to waste quite less memory \nis a good indicator already.\n\nI did modify the progress display to show accounted memory that was \nallocated vs memory that was freed but still not released to the system.  \nAt least that gives you an idea of memory allocation and fragmentation \nwith glibc in real time:\n\ndiff --git a/progress.c b/progress.c\nindex d19f80c..46ac9ef 100644\n--- a/progress.c\n+++ b/progress.c\n@@ -8,6 +8,7 @@\n  * published by the Free Software Foundation.\n  */\n \n+#include <malloc.h>\n #include \"git-compat-util.h\"\n #include \"progress.h\"\n \n@@ -94,10 +95,12 @@ static int display(struct progress *progress, unsigned n, const char *done)\n \tif (progress->total) {\n \t\tunsigned percent = n * 100 / progress->total;\n \t\tif (percent != progress->last_percent || progress_update) {\n+\t\t\tstruct mallinfo m = mallinfo();\n \t\t\tprogress->last_percent = percent;\n-\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u)%s%s\",\n-\t\t\t\tprogress->title, percent, n,\n-\t\t\t\tprogress->total, tp, eol);\n+\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u) %u/%uMB%s%s\",\n+\t\t\t\tprogress->title, percent, n, progress->total,\n+\t\t\t\tm.uordblks >> 18, m.fordblks >> 18,\n+\t\t\t\ttp, eol);\n \t\t\tfflush(stderr);\n \t\t\tprogress_update = 0;\n \t\t\treturn 1;\n\nThis shows that at some point the repack goes into a big memory surge.  \nI don't have enough RAM to see how fragmented memory gets though, since \nit starts swapping around 50% done with 2 threads.\n\nWith only 1 thread, memory usage grows significantly at around 11% with \na pretty noticeable slowdown in the progress rate.\n\nSo I think the theory goes like this:\n\nThere is a block of big objects together in the list somewhere.  \nInitially, all those big objects are assigned to thread #1 out of 4.  \nBecause those objects are big, they get really slow to delta compress, \nand storing them all in a window with 250 slots takes significant \nmemory.\n\nThreads 2, 3, and 4 have \"easy\" work loads, so they complete fairly \nquicly compared to thread #1.  But since the progress display is global \nthen you won't notice that one thread is actually crawling slowly.\n\nTo keep all threads busy until the end, those threads that are done with \ntheir work load will steal some work from another thread, choosing the \none with the largest remaining work.  That is most likely thread #1.  So \nas threads 2, 3, and 4 complete, they will steal from thread 1 and \npopulate their own window with those big objects too, and get slow too.\n\nAnd because all threads gets to work on those big objects towards the \nend, the progress display will then show a significant slowdown, and \nmemory usage will almost quadruple.\n\nAdd memory fragmentation to that and you have a clogged system.\n\nSolution: \n\n\tpack.deltacachesize=1\n\tpack.windowmemory=16M\n\nLimiting the window memory to 16MB will automatically shrink the window \nsize when big objects are encountered, therefore keeping much fewer of \nthose objects at the same time in memory, which in turn means they will \nbe processed much more quickly.  And somehow that must help with memory \nfragmentation as well.\n\nSetting pack.deltacachesize to 1 is simply to disable the caching of \ndelta results entirely which will only slow down the writing phase, but \nI wanted to keep it out of the picture for now.\n\nWith the above settings, I'm currently repacking the gcc repo with 2 \nthreads, and memory allocation never exceeded 700m virt and 400m res, \nwhile the mallinfo shows about 350MB, and progress has reached 90% which \nhas never occurred on this machine with the 300MB source pack so far.\n\n\nNicolas\n"},{"id":"62868","messageId":"85d4tc8hi8.fsf@lola.goethe.zz","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712112057390.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-12-12T08:05:51Z","receivedAt":"2007-12-12T08:05:51Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> Well... This is weird.\n>\n> It seems that memory fragmentation is really really killing us here.  \n> The fact that the Google allocator did manage to waste quite less memory \n> is a good indicator already.\n\nMaybe an malloc/free/mmap wrapper that records the requested sizes and\nalloc/free order and dumps them to file so that one can make a compact\ngit-free standalone test case for the glibc maintainers might be a good\nthing.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"62915","messageId":"alpine.LFD.0.99999.0712120743040.555@xanadu.home","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712112057390.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-12T15:48:12Z","receivedAt":"2007-12-12T15:48:12Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 12 Dec 2007, Nicolas Pitre wrote:\n\n> Add memory fragmentation to that and you have a clogged system.\n> \n> Solution: \n> \n> \tpack.deltacachesize=1\n> \tpack.windowmemory=16M\n> \n> Limiting the window memory to 16MB will automatically shrink the window \n> size when big objects are encountered, therefore keeping much fewer of \n> those objects at the same time in memory, which in turn means they will \n> be processed much more quickly.  And somehow that must help with memory \n> fragmentation as well.\n\nOK scrap that.\n\nWhen I returned to the computer this morning, the repack was \ncompleted... with a 1.3GB pack instead.\n\nSo... The gcc repo apparently really needs a large window to efficiently \ncompress those large objects.\n\nBut when those large objects are already well deltified and you repack \nagain with a large window, somehow the memory allocator is way more \ninvolved, probably even \nmore so when there are several threads in parallel amplifying the issue, \nand things probably get to a point of no return with regard to memory \nfragmentation after a while.\n\nSo... my conclusion is that the glibc allocator has fragmentation issues \nwith this work load, given the notable difference with the Google \nallocator, which itself might not be completely immune to fragmentation \nissues of its own.  And because the gcc repo requires a large window of \nbig objects to get good compression, then you're better not using 4 \nthreads to repack it with -a -f.  The fact that the size of the source \npack has such an influence is probably only because the increased usage \nof the delta base object cache is playing a role in the global memory \nallocation pattern, allowing for the bad fragmentation issue to occur.\n\nIf you could run one last test with the mallinfo patch I posted, without \nthe pack.windowmemory setting, and adding the reported values along with \nthose from top, then we could formally conclude to memory fragmentation \nissues.\n\nSo I don't think Git itself is actually bad.  The gcc repo most \ncertainly constitute a nasty use case for memory allocators, but I don't \nthink there is much we can do about it besides possibly implementing our \nown memory allocator with active defragmentation where possible (read \nmemcpy) at some point to give glibc's allocator some chance to breathe a \nbit more.\n\nIn the mean time you might have to use only one thread and lots of \nmemory to repack the gcc repo, or find the perfect memory allocator to \nbe used with Git.  After all, packing the whole gcc history to around \n230MB is quite a stunt but it requires sufficient resources to \nachieve it. Fortunately, like Linus said, such a wholesale repack is not \nsomething that most users have to do anyway.\n\n\nNicolas\n"},{"id":"62923","messageId":"alpine.LFD.0.99999.0712121106400.555@xanadu.home","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712112057390.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-12T16:13:52Z","receivedAt":"2007-12-12T16:13:52Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 12 Dec 2007, Nicolas Pitre wrote:\n\n> I did modify the progress display to show accounted memory that was \n> allocated vs memory that was freed but still not released to the system.  \n> At least that gives you an idea of memory allocation and fragmentation \n> with glibc in real time:\n> \n> diff --git a/progress.c b/progress.c\n> index d19f80c..46ac9ef 100644\n> --- a/progress.c\n> +++ b/progress.c\n> @@ -8,6 +8,7 @@\n>   * published by the Free Software Foundation.\n>   */\n>  \n> +#include <malloc.h>\n>  #include \"git-compat-util.h\"\n>  #include \"progress.h\"\n>  \n> @@ -94,10 +95,12 @@ static int display(struct progress *progress, unsigned n, const char *done)\n>  \tif (progress->total) {\n>  \t\tunsigned percent = n * 100 / progress->total;\n>  \t\tif (percent != progress->last_percent || progress_update) {\n> +\t\t\tstruct mallinfo m = mallinfo();\n>  \t\t\tprogress->last_percent = percent;\n> -\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u)%s%s\",\n> -\t\t\t\tprogress->title, percent, n,\n> -\t\t\t\tprogress->total, tp, eol);\n> +\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u) %u/%uMB%s%s\",\n> +\t\t\t\tprogress->title, percent, n, progress->total,\n> +\t\t\t\tm.uordblks >> 18, m.fordblks >> 18,\n> +\t\t\t\ttp, eol);\n\nNote: I didn't know what unit of memory those blocks represents, so the \nshift is most probably wrong.\n\n\nNicolas\n"},{"id":"62925","messageId":"fjp1iu$rv3$1@ger.gmane.org","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712120743040.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2007-12-12T16:17:43Z","receivedAt":"2007-12-12T16:17:43Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> When I returned to the computer this morning, the repack was \n> completed... with a 1.3GB pack instead.\n> \n> So... The gcc repo apparently really needs a large window to efficiently \n> compress those large objects.\n\nSo, am I right that if you have a very well-done pack (such as gcc's), \nyou might want to repack in two phases:\n\n- first discarding the old deltas and using a small window, thus \nproducing a bad pack that can be repacked without humongous amounts of \nmemory...\n\n- ... then discarding the old deltas and producing another \nwell-compressed pack?\n\nPaolo\n"},{"id":"62933","messageId":"alpine.LFD.0.9999.0712120826440.25032@woody.linux-foundation.org","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712120743040.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T16:37:10Z","receivedAt":"2007-12-12T16:37:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 12 Dec 2007, Nicolas Pitre wrote:\n> \n> So... my conclusion is that the glibc allocator has fragmentation issues \n> with this work load, given the notable difference with the Google \n> allocator, which itself might not be completely immune to fragmentation \n> issues of its own. \n\nYes.\n\nNote that delta following involves patterns something like\n\n   allocate (small) space for delta\n   for i in (1..depth) {\n\tallocate large space for base\n\tallocate large space for result\n\t.. apply delta ..\n\tfree large space for base\n\tfree small space for delta\n   }\n\nso if you have some stupid heap algorithm that doesn't try to merge and \nre-use free'd spaces very aggressively (because that takes CPU time!), you \nmight have memory usage be horribly inflated by the heap having all those \nholes for all the objects that got free'd in the chain that don't get \naggressively re-used.\n\nThreaded memory allocators then make this worse by probably using totally \ndifferent heaps for different threads (in order to avoid locking), so they \nwill *all* have the fragmentation issue.\n\nAnd if you *really* want to cause trouble for a memory allocator, what you \nshould try to do is to allocate the memory in one thread, and free it in \nanother, and then things can really explode (the freeing thread notices \nthat the allocation is not in its thread-local heap, so instead of really \nfreeing it, it puts it on a separate list of areas to be freed later by \nthe original thread when it needs memory - or worse, it adds it to the \nlocal thread list, and makes it effectively totally impossible to then \never merge different free'd allocations ever again because the freed \nthings will be on different heap lists!).\n\nI'm not saying that particular case happens in git, I'm just saying that \nit's not unheard of. And with the delta cache and the object lookup, it's \nnot at _all_ impossible that we hit the \"allocate in one thread, free in \nanother\" case!\n\n\t\tLinus\n"},{"id":"62934","messageId":"20071212.084212.02518392.davem@davemloft.net","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712120826440.25032@woody.linux-foundation.org","subject":"Re: Something is broken in repack","fromName":"David Miller","fromEmail":"davem@davemloft.net","sentAt":"2007-12-12T16:42:12Z","receivedAt":"2007-12-12T16:42:12Z","isPatch":false,"sender":{"key":"davem@davemloft.net","avatar":null},"body":"From: Linus Torvalds <torvalds@linux-foundation.org>\nDate: Wed, 12 Dec 2007 08:37:10 -0800 (PST)\n\n> I'm not saying that particular case happens in git, I'm just saying that \n> it's not unheard of. And with the delta cache and the object lookup, it's \n> not at _all_ impossible that we hit the \"allocate in one thread, free in \n> another\" case!\n\nOne thing that supports these theories is that, while running\nthese large repacks, I notice that the RSS is roughly 2/3 of\nthe amount of virtual address space allocated.\n\nI personally don't think it's unreasonable for GIT to have it's\nown customized allocator at least for certain object types.\n"},{"id":"62935","messageId":"alpine.LFD.0.9999.0712120848130.25032@woody.linux-foundation.org","threadId":"11182","inReplyTo":"20071212.084212.02518392.davem@davemloft.net","subject":"Re: Something is broken in repack","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-12-12T16:54:28Z","receivedAt":"2007-12-12T16:54:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 12 Dec 2007, David Miller wrote:\n> \n> I personally don't think it's unreasonable for GIT to have it's\n> own customized allocator at least for certain object types.\n\nWell, we actually already *do* have a customized allocator, but currently \nonly for the actual core \"object descriptor\" that really just has the SHA1 \nand object flags in it (and a few extra words depending on object type).\n\nThose are critical for certain loads, and small too (so using the standard \nallocator wasted a _lot_ of memory). In addition, they're fixed-size and \nnever free'd, so a specialized allocator really can do a lot better than \nany general-purpose memory allocator ever could.\n\nBut the actual object *contents* are currently all allocated with whatever \nthe standard libc malloc/free allocator is that you compile for (or load \ndynamically). Havign a specialized allocator for them is a much more \ninvolved issue, exactly because we do have interesting allocation patterns \netc.\n\nThat said, at least those object allocations are all single-threaded (for \nright now, at least), so even when git does multi-threaded stuff, the core \nsha1_file.c stuff is always run under a single lock, and a simpler \nallocator that doesn't care about threads is likely to be much better than \none that tries to have thread-local heaps etc.\n\nI suspect that is what the google allocator does. It probably doesn't have \nper-thread heaps, it just uses locking (and quite possibly things like \nper-*size* heaps, which is much more memory-efficient and helps avoid some \nof the fragmentation problems). \n\nLocking is much slower than per-thread accesses, but it doesn't have the \nissues with per-thread-fragmentation and all the problems with one thread \nallocating and another one freeing.\n\n\t\t\tLinus\n"},{"id":"62936","messageId":"9e4733910712120912l342350f2i1f190c45730108f2@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712120826440.25032@woody.linux-foundation.org","subject":"Re: Something is broken in repack","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-12-12T17:12:04Z","receivedAt":"2007-12-12T17:12:04Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 12/12/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Wed, 12 Dec 2007, Nicolas Pitre wrote:\n> >\n> > So... my conclusion is that the glibc allocator has fragmentation issues\n> > with this work load, given the notable difference with the Google\n> > allocator, which itself might not be completely immune to fragmentation\n> > issues of its own.\n>\n> Yes.\n>\n> Note that delta following involves patterns something like\n>\n>    allocate (small) space for delta\n>    for i in (1..depth) {\n>         allocate large space for base\n>         allocate large space for result\n>         .. apply delta ..\n>         free large space for base\n>         free small space for delta\n>    }\n\nIs it hard to hack up something that statically allocates a big block\nof memory per thread for these two and then just reuses it?\n   allocate (small) space for delta\n   allocate large space for base\n\nThe alternating between long term and short term allocations\ndefinitely aggravates fragmentation.\n\n>\n> so if you have some stupid heap algorithm that doesn't try to merge and\n> re-use free'd spaces very aggressively (because that takes CPU time!), you\n> might have memory usage be horribly inflated by the heap having all those\n> holes for all the objects that got free'd in the chain that don't get\n> aggressively re-used.\n>\n> Threaded memory allocators then make this worse by probably using totally\n> different heaps for different threads (in order to avoid locking), so they\n> will *all* have the fragmentation issue.\n>\n> And if you *really* want to cause trouble for a memory allocator, what you\n> should try to do is to allocate the memory in one thread, and free it in\n> another, and then things can really explode (the freeing thread notices\n> that the allocation is not in its thread-local heap, so instead of really\n> freeing it, it puts it on a separate list of areas to be freed later by\n> the original thread when it needs memory - or worse, it adds it to the\n> local thread list, and makes it effectively totally impossible to then\n> ever merge different free'd allocations ever again because the freed\n> things will be on different heap lists!).\n>\n> I'm not saying that particular case happens in git, I'm just saying that\n> it's not unheard of. And with the delta cache and the object lookup, it's\n> not at _all_ impossible that we hit the \"allocate in one thread, free in\n> another\" case!\n>\n>                 Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"62987","messageId":"4760E005.6040102@op5.se","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712121106400.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2007-12-13T07:32:21Z","receivedAt":"2007-12-13T07:32:21Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Nicolas Pitre wrote:\n> On Wed, 12 Dec 2007, Nicolas Pitre wrote:\n> \n>> I did modify the progress display to show accounted memory that was \n>> allocated vs memory that was freed but still not released to the system.  \n>> At least that gives you an idea of memory allocation and fragmentation \n>> with glibc in real time:\n>>\n>> diff --git a/progress.c b/progress.c\n>> index d19f80c..46ac9ef 100644\n>> --- a/progress.c\n>> +++ b/progress.c\n>> @@ -8,6 +8,7 @@\n>>   * published by the Free Software Foundation.\n>>   */\n>>  \n>> +#include <malloc.h>\n>>  #include \"git-compat-util.h\"\n>>  #include \"progress.h\"\n>>  \n>> @@ -94,10 +95,12 @@ static int display(struct progress *progress, unsigned n, const char *done)\n>>  \tif (progress->total) {\n>>  \t\tunsigned percent = n * 100 / progress->total;\n>>  \t\tif (percent != progress->last_percent || progress_update) {\n>> +\t\t\tstruct mallinfo m = mallinfo();\n>>  \t\t\tprogress->last_percent = percent;\n>> -\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u)%s%s\",\n>> -\t\t\t\tprogress->title, percent, n,\n>> -\t\t\t\tprogress->total, tp, eol);\n>> +\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u) %u/%uMB%s%s\",\n>> +\t\t\t\tprogress->title, percent, n, progress->total,\n>> +\t\t\t\tm.uordblks >> 18, m.fordblks >> 18,\n>> +\t\t\t\ttp, eol);\n> \n> Note: I didn't know what unit of memory those blocks represents, so the \n> shift is most probably wrong.\n> \n\nMe neither, but it appears to me as if hblkhd holds the actual memory\nconsumed by the process. It seems to store the information in bytes,\nwhich I find a bit dubious unless glibc has some internal multiplier.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"63005","messageId":"fcaeb9bf0712130532s79aa7afeve6f018f9430ab3b3@mail.gmail.com","threadId":"11182","inReplyTo":"alpine.LFD.0.99999.0712120743040.555@xanadu.home","subject":"Re: Something is broken in repack","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2007-12-13T13:32:02Z","receivedAt":"2007-12-13T13:32:02Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Dec 12, 2007 10:48 PM, Nicolas Pitre <nico@cam.org> wrote:\n> In the mean time you might have to use only one thread and lots of\n> memory to repack the gcc repo, or find the perfect memory allocator to\n> be used with Git.  After all, packing the whole gcc history to around\n> 230MB is quite a stunt but it requires sufficient resources to\n> achieve it. Fortunately, like Linus said, such a wholesale repack is not\n> something that most users have to do anyway.\n\nIs there an alternative to \"git repack -a -d\" that repacks everything\nbut the first pack?\n-- \nDuy\n"},{"id":"63023","messageId":"fjrj9k$n6k$1@ger.gmane.org","threadId":"11182","inReplyTo":"fcaeb9bf0712130532s79aa7afeve6f018f9430ab3b3@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2007-12-13T15:32:03Z","receivedAt":"2007-12-13T15:32:03Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"Nguyen Thai Ngoc Duy wrote:\n> On Dec 12, 2007 10:48 PM, Nicolas Pitre <nico@cam.org> wrote:\n>> In the mean time you might have to use only one thread and lots of\n>> memory to repack the gcc repo, or find the perfect memory allocator to\n>> be used with Git.  After all, packing the whole gcc history to around\n>> 230MB is quite a stunt but it requires sufficient resources to\n>> achieve it. Fortunately, like Linus said, such a wholesale repack is not\n>> something that most users have to do anyway.\n> \n> Is there an alternative to \"git repack -a -d\" that repacks everything\n> but the first pack?\n\nThat would be a pretty good idea for big repositories.  If I were to \nimplement it, I would actually add a .git/config option like \npack.permanent so that more than one pack could be made permanent; then \nto repack really really everything you'd need \"git repack -a -a -d\".\n\nPaolo\n"},{"id":"63032","messageId":"47615E04.8000400@gnu.org","threadId":"11182","inReplyTo":"fjrj9k$n6k$1@ger.gmane.org","subject":"Re: Something is broken in repack","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2007-12-13T16:29:56Z","receivedAt":"2007-12-13T16:29:56Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n>> Is there an alternative to \"git repack -a -d\" that repacks everything\n>> but the first pack?\n> \n> That would be a pretty good idea for big repositories.  If I were to \n> implement it, I would actually add a .git/config option like \n> pack.permanent so that more than one pack could be made permanent; then \n> to repack really really everything you'd need \"git repack -a -a -d\".\n\nActually there is something like this, as seen from the source of \ngit-repack:\n\n             for e in `cd \"$PACKDIR\" && find . -type f -name '*.pack' \\\n                      | sed -e 's/^\\.\\///' -e 's/\\.pack$//'`\n             do\n                     if [ -e \"$PACKDIR/$e.keep\" ]; then\n                             : keep\n                     else\n                             args=\"$args --unpacked=$e.pack\"\n                             existing=\"$existing $e\"\n                     fi\n             done\n\nSo, just create a file named as the pack, but with extension \".keep\".\n\nPaolo\n"},{"id":"63035","messageId":"47616044.7070504@viscovery.net","threadId":"11182","inReplyTo":"fjrj9k$n6k$1@ger.gmane.org","subject":"Re: Something is broken in repack","fromName":"Johannes Sixt","fromEmail":"j.sixt@viscovery.net","sentAt":"2007-12-13T16:39:32Z","receivedAt":"2007-12-13T16:39:32Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Paolo Bonzini schrieb:\n> Nguyen Thai Ngoc Duy wrote:\n>> On Dec 12, 2007 10:48 PM, Nicolas Pitre <nico@cam.org> wrote:\n>>> In the mean time you might have to use only one thread and lots of\n>>> memory to repack the gcc repo, or find the perfect memory allocator to\n>>> be used with Git.  After all, packing the whole gcc history to around\n>>> 230MB is quite a stunt but it requires sufficient resources to\n>>> achieve it. Fortunately, like Linus said, such a wholesale repack is not\n>>> something that most users have to do anyway.\n>>\n>> Is there an alternative to \"git repack -a -d\" that repacks everything\n>> but the first pack?\n> \n> That would be a pretty good idea for big repositories.  If I were to\n> implement it, I would actually add a .git/config option like\n> pack.permanent so that more than one pack could be made permanent; then\n> to repack really really everything you'd need \"git repack -a -a -d\".\n\nIt's already there: If you have a pack .git/objects/pack/pack-foo.pack, then\n\"touch .git/objects/pack/pack-foo.keep\" marks the pack as precious.\n\n-- Hannes\n"},{"id":"63068","messageId":"fjskqt$eap$1@ger.gmane.org","threadId":"11182","inReplyTo":"47616044.7070504@viscovery.net","subject":"Re: Something is broken in repack","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-14T01:04:29Z","receivedAt":"2007-12-14T01:04:29Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Sixt wrote:\n> Paolo Bonzini schrieb:\n>> Nguyen Thai Ngoc Duy wrote:\n>>>\n>>> Is there an alternative to \"git repack -a -d\" that repacks everything\n>>> but the first pack?\n>> \n>> That would be a pretty good idea for big repositories.  If I were to\n>> implement it, I would actually add a .git/config option like\n>> pack.permanent so that more than one pack could be made permanent; then\n>> to repack really really everything you'd need \"git repack -a -a -d\".\n> \n> It's already there: If you have a pack .git/objects/pack/pack-foo.pack, then\n> \"touch .git/objects/pack/pack-foo.keep\" marks the pack as precious.\n\nActually you can (and probably should) put the one line with the _reason_\npack is to be kept in the *.keep file.\n\nHmmm... it is even documented in git-gc(1)... and git-index-pack(1) of\nall things.\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"63094","messageId":"fjt6vm$n7d$1@ger.gmane.org","threadId":"11182","inReplyTo":"fjskqt$eap$1@ger.gmane.org","subject":"Re: Something is broken in repack","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2007-12-14T06:14:14Z","receivedAt":"2007-12-14T06:14:14Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"> Hmmm... it is even documented in git-gc(1)... and git-index-pack(1) of\n> all things.\n\nI found that the .keep file is not transmitted over the network (at \nleast I tried with git+ssh:// and http:// protocols), however.\n\nPaolo\n"},{"id":"63095","messageId":"fcaeb9bf0712132224u54ca845ap4836dfe1cda37b29@mail.gmail.com","threadId":"11182","inReplyTo":"fjt6vm$n7d$1@ger.gmane.org","subject":"Re: Something is broken in repack","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2007-12-14T06:24:07Z","receivedAt":"2007-12-14T06:24:07Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Dec 14, 2007 1:14 PM, Paolo Bonzini <bonzini@gnu.org> wrote:\n> > Hmmm... it is even documented in git-gc(1)... and git-index-pack(1) of\n> > all things.\n>\n> I found that the .keep file is not transmitted over the network (at\n> least I tried with git+ssh:// and http:// protocols), however.\n\nI'm thinking about \"git clone --keep\" to mark initial packs precious.\nBut 'git clone' is under rewrite to C. Let's wait until C rewrite is\ndone.\n-- \nDuy\n"},{"id":"63127","messageId":"fjtect$8qn$1@ger.gmane.org","threadId":"11182","inReplyTo":"fcaeb9bf0712132224u54ca845ap4836dfe1cda37b29@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2007-12-14T08:20:45Z","receivedAt":"2007-12-14T08:20:45Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> I'm thinking about \"git clone --keep\" to mark initial packs precious.\n> But 'git clone' is under rewrite to C. Let's wait until C rewrite is\n> done.\n\nIt should be the default, IMHO.\n\nPaolo\n"},{"id":"63130","messageId":"1197622912.898.53.camel@brick","threadId":"11182","inReplyTo":"fjtect$8qn$1@ger.gmane.org","subject":"Re: Something is broken in repack","fromName":"Harvey Harrison","fromEmail":"harvey.harrison@gmail.com","sentAt":"2007-12-14T09:01:52Z","receivedAt":"2007-12-14T09:01:52Z","isPatch":false,"sender":{"key":"harvey.harrison@gmail.com","avatar":null},"body":"On Fri, 2007-12-14 at 09:20 +0100, Paolo Bonzini wrote:\n> > I'm thinking about \"git clone --keep\" to mark initial packs precious.\n> > But 'git clone' is under rewrite to C. Let's wait until C rewrite is\n> > done.\n> \n> It should be the default, IMHO.\n> \n\nWhile it doesn't mark the packs as .keep, git will reuse all of the old\ndeltas you got in the original clone, so you're not losing anything.\n\nHarvey\n"},{"id":"63135","messageId":"m3odctr245.fsf@roke.D-201","threadId":"11182","inReplyTo":"fcaeb9bf0712132224u54ca845ap4836dfe1cda37b29@mail.gmail.com","subject":"Re: Something is broken in repack","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-14T10:40:25Z","receivedAt":"2007-12-14T10:40:25Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Nguyen Thai Ngoc Duy\" <pclouds@gmail.com> writes:\n\n> On Dec 14, 2007 1:14 PM, Paolo Bonzini <bonzini@gnu.org> wrote:\n> > > Hmmm... it is even documented in git-gc(1)... and git-index-pack(1) of\n> > > all things.\n> >\n> > I found that the .keep file is not transmitted over the network (at\n> > least I tried with git+ssh:// and http:// protocols), however.\n> \n> I'm thinking about \"git clone --keep\" to mark initial packs precious.\n> But 'git clone' is under rewrite to C. Let's wait until C rewrite is\n> done.\n\nBut if you clone via network, pack might be network optimized if you\nuse \"smart\" transport, not disk optimized, at least with current git\nwhich regenerates pack also on clone AFAIK.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"63136","messageId":"fcaeb9bf0712140252j42a3cff0r33b889e355d41dd@mail.gmail.com","threadId":"11182","inReplyTo":"m3odctr245.fsf@roke.D-201","subject":"Re: Something is broken in repack","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2007-12-14T10:52:05Z","receivedAt":"2007-12-14T10:52:05Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Dec 14, 2007 4:01 PM, Harvey Harrison <harvey.harrison@gmail.com> wrote:\n> While it doesn't mark the packs as .keep, git will reuse all of the old\n> deltas you got in the original clone, so you're not losing anything.\n\nThere is another reason I want it. I have an ~800MB pack and I don't\nwant git to rewrite  the pack every time I repack my changes. So it's\nkind of disk-wise (don't require 800MB on disk to prepare new pack,\nand don't write too much).\n\nOn Dec 14, 2007 5:40 PM, Jakub Narebski <jnareb@gmail.com> wrote:\n> But if you clone via network, pack might be network optimized if you\n> use \"smart\" transport, not disk optimized, at least with current git\n> which regenerates pack also on clone AFAIK.\n\nUm.. that's ok it just regenerate once.\n\n-- \nDuy\n"},{"id":"63163","messageId":"alpine.LFD.0.999999.0712140823580.8467@xanadu.home","threadId":"11182","inReplyTo":"fjt6vm$n7d$1@ger.gmane.org","subject":"Re: Something is broken in repack","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-12-14T13:25:16Z","receivedAt":"2007-12-14T13:25:16Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 14 Dec 2007, Paolo Bonzini wrote:\n\n> > Hmmm... it is even documented in git-gc(1)... and git-index-pack(1) of\n> > all things.\n> \n> I found that the .keep file is not transmitted over the network (at least I\n> tried with git+ssh:// and http:// protocols), however.\n\nThat is a local policy.\n\n\nNicolas\n"},{"id":"63174","messageId":"20071214160326.2424.qmail@md.dent.med.uni-muenchen.de","threadId":"11182","inReplyTo":"4760E005.6040102@op5.se","subject":"Re: Something is broken in repack","fromName":"Wolfram Gloger","fromEmail":"wmglo@dent.med.uni-muenchen.de","sentAt":null,"receivedAt":"2007-12-14T16:01:42Z","isPatch":false,"sender":{"key":"wmglo@dent.med.uni-muenchen.de","avatar":null},"body":"Hi,\n\n> >>  \tif (progress->total) {\n> >>  \t\tunsigned percent = n * 100 / progress->total;\n> >>  \t\tif (percent != progress->last_percent || progress_update) {\n> >> +\t\t\tstruct mallinfo m = mallinfo();\n> >>  \t\t\tprogress->last_percent = percent;\n> >> -\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u)%s%s\",\n> >> -\t\t\t\tprogress->title, percent, n,\n> >> -\t\t\t\tprogress->total, tp, eol);\n> >> +\t\t\tfprintf(stderr, \"%s: %3u%% (%u/%u) %u/%uMB%s%s\",\n> >> +\t\t\t\tprogress->title, percent, n, progress->total,\n> >> +\t\t\t\tm.uordblks >> 18, m.fordblks >> 18,\n> >> +\t\t\t\ttp, eol);\n> > \n> > Note: I didn't know what unit of memory those blocks represents, so the \n> > shift is most probably wrong.\n> > \n> \n> Me neither, but it appears to me as if hblkhd holds the actual memory\n> consumed by the process. It seems to store the information in bytes,\n> which I find a bit dubious unless glibc has some internal multiplier.\n\nmallinfo() will only give you the used memory for the main arena.\nWhen you have separate arenas (likely when concurrent threads have\nbeen used), the only way to get the full picture is to call\nmalloc_stats(), which prints to stderr.\n\nRegards,\nWolfram.\n"},{"id":"63175","messageId":"20071214161236.3080.qmail@md.dent.med.uni-muenchen.de","threadId":"11182","inReplyTo":"alpine.LFD.0.9999.0712120826440.25032@woody.linux-foundation.org","subject":"Re: Something is broken in repack","fromName":"Wolfram Gloger","fromEmail":"wmglo@dent.med.uni-muenchen.de","sentAt":null,"receivedAt":"2007-12-14T16:01:42Z","isPatch":false,"sender":{"key":"wmglo@dent.med.uni-muenchen.de","avatar":null},"body":"Hi,\n\n> Note that delta following involves patterns something like\n> \n>    allocate (small) space for delta\n>    for i in (1..depth) {\n> \tallocate large space for base\n> \tallocate large space for result\n> \t.. apply delta ..\n> \tfree large space for base\n> \tfree small space for delta\n>    }\n> \n> so if you have some stupid heap algorithm that doesn't try to merge and \n> re-use free'd spaces very aggressively (because that takes CPU time!),\n\nptmalloc2 (in glibc) _per arena_ is basically best-fit.  This is the\nbest known general strategy, but it certainly cannot be the best in\nevery case.\n\n> you \n> might have memory usage be horribly inflated by the heap having all those \n> holes for all the objects that got free'd in the chain that don't get \n> aggressively re-used.\n\nIt depends how large 'large' is -- if it exceeds the mmap() threshold\n(settable with mallopt(M_MMAP_THRESHOLD, ...))\nthe 'large' spaces will be allocated with mmap() and won't cause\nany internal fragmentation.\nIt might pay to experiment with this parameter if it is hard to\navoid the alloc/free large space sequence.\n\n> Threaded memory allocators then make this worse by probably using totally \n> different heaps for different threads (in order to avoid locking), so they \n> will *all* have the fragmentation issue.\n\nIndeed.\n\nCould someone perhaps try ptmalloc3\n(http://malloc.de/malloc/ptmalloc3-current.tar.gz) on this case?\n\nThanks,\nWolfram.\n"},{"id":"63176","messageId":"20071214161858.3506.qmail@md.dent.med.uni-muenchen.de","threadId":"11182","inReplyTo":"85d4tc8hi8.fsf@lola.goethe.zz","subject":"Re: Something is broken in repack","fromName":"Wolfram Gloger","fromEmail":"wmglo@dent.med.uni-muenchen.de","sentAt":null,"receivedAt":"2007-12-14T16:01:42Z","isPatch":false,"sender":{"key":"wmglo@dent.med.uni-muenchen.de","avatar":null},"body":"Hi,\n\n> Maybe an malloc/free/mmap wrapper that records the requested sizes and\n> alloc/free order and dumps them to file so that one can make a compact\n> git-free standalone test case for the glibc maintainers might be a good\n> thing.\n\nI already have such a wrapper:\n\nhttp://malloc.de/malloc/mtrace-20060529.tar.gz\n\nBut note that it does interfere with the thread scheduling, so it\ncan't record the exact same allocation pattern as when not using the\nwrapper.\n\nRegards,\nWolfram.\n"},{"id":"63179","messageId":"85r6hptecs.fsf@lola.goethe.zz","threadId":"11182","inReplyTo":"20071214161236.3080.qmail@md.dent.med.uni-muenchen.de","subject":"Re: Something is broken in repack","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-12-14T16:45:07Z","receivedAt":"2007-12-14T16:45:07Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Wolfram Gloger <wmglo@dent.med.uni-muenchen.de> writes:\n\n> Hi,\n>\n>> Note that delta following involves patterns something like\n>> \n>>    allocate (small) space for delta\n>>    for i in (1..depth) {\n>> \tallocate large space for base\n>> \tallocate large space for result\n>> \t.. apply delta ..\n>> \tfree large space for base\n>> \tfree small space for delta\n>>    }\n>> \n>> so if you have some stupid heap algorithm that doesn't try to merge and \n>> re-use free'd spaces very aggressively (because that takes CPU time!),\n>\n> ptmalloc2 (in glibc) _per arena_ is basically best-fit.  This is the\n> best known general strategy,\n\nUh what?  Someone crank out his copy of \"The Art of Computer\nProgramming\", I think volume 1.  Best fit is known (analyzed and proven\nand documented decades ago) to be one of the worst strategies for memory\nallocation.  Exactly because it leads to huge fragmentation problems.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"63180","messageId":"20071214165937.6405.qmail@md.dent.med.uni-muenchen.de","threadId":"11182","inReplyTo":"85r6hptecs.fsf@lola.goethe.zz","subject":"Re: Something is broken in repack","fromName":"Wolfram Gloger","fromEmail":"wmglo@dent.med.uni-muenchen.de","sentAt":null,"receivedAt":"2007-12-14T16:45:07Z","isPatch":false,"sender":{"key":"wmglo@dent.med.uni-muenchen.de","avatar":null},"body":"Hi,\n\n> Uh what?  Someone crank out his copy of \"The Art of Computer\n> Programming\", I think volume 1.  Best fit is known (analyzed and proven\n> and documented decades ago) to be one of the worst strategies for memory\n> allocation.  Exactly because it leads to huge fragmentation problems.\n\nWell, quoting http://gee.cs.oswego.edu/dl/html/malloc.html:\n\n\"As shown by Wilson et al, best-fit schemes (of various kinds and\napproximations) tend to produce the least fragmentation on real loads\ncompared to other general approaches such as first-fit.\"\n\nSee [Wilson 1995] ftp://ftp.cs.utexas.edu/pub/garbage/allocsrv.ps for\nmore details and references.\n\nRegards,\nWolfram.\n"}]}