{"thread":{"id":"29693","subject":"[FYI] very large text files and their problems.","startedAt":"2012-02-22T15:49:26Z","lastAt":"2012-02-24T12:55:52Z","messageCount":5,"participants":["Ian Kumlien","Nguyen Thai Ngoc Duy"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"185180","messageId":"20120222154926.GC11202@pomac.netswarm.net","threadId":"29693","inReplyTo":null,"subject":"[FYI] very large text files and their problems.","fromName":"Ian Kumlien","fromEmail":"pomac@vapor.com","sentAt":"2012-02-22T15:49:26Z","receivedAt":"2012-02-22T15:49:26Z","isPatch":false,"sender":{"key":"pomac@vapor.com","avatar":null},"body":"Hi, \n\nWe just saw a interesting issue, git compressed a ~3.4 gb project to ~57\nmb. But when we tried to clone it on a big machine we got:\n\nfatal: Out of memory, malloc failed (tried to allocate\n18446744072724798634 bytes)\n\nThis is already fixed in the 1.7.10 mainline - but it also seems like\ngit needs to have atleast the same ammount of memory as the largest\nfile free... Couldn't this be worked around?\n\nOn a (32 bit) machine with 4GB memory - results in:\nfatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)\n\n(and i see how this could be a problem, but couldn't it be mitigated? or\nis it bydesign and intended behaviour?)\n\nI'm not subscribed to please keep me in CC.\n\n/Ian Kumlien\n"},{"id":"185181","messageId":"CACsJy8Bdbegs7QdztvsFnKPcpAX5UL7s7uc37wF3_nF4kJQjrQ@mail.gmail.com","threadId":"29693","inReplyTo":"20120222154926.GC11202@pomac.netswarm.net","subject":"Re: [FYI] very large text files and their problems.","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-22T16:18:19Z","receivedAt":"2012-02-22T16:18:19Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Feb 22, 2012 at 10:49 PM, Ian Kumlien <pomac@vapor.com> wrote:\n> Hi,\n>\n> We just saw a interesting issue, git compressed a ~3.4 gb project to ~57 mb.\n\nHow big are those files? How many of them? How often do they change?\n\n> But when we tried to clone it on a big machine we got:\n>\n> fatal: Out of memory, malloc failed (tried to allocate\n> 18446744072724798634 bytes)\n>\n> This is already fixed in the 1.7.10 mainline - but it also seems like\n\nDoes 1.7.9 have this problem?\n\n> git needs to have atleast the same ammount of memory as the largest\n> file free... Couldn't this be worked around?\n>\n> On a (32 bit) machine with 4GB memory - results in:\n> fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)\n>\n> (and i see how this could be a problem, but couldn't it be mitigated? or\n> is it bydesign and intended behaviour?)\n\nI think that it's delta resolving that hogs all your memory. If your\nfiles are smaller than 512M, try lower core.bigFileThreshold. The\ntopic jc/split-blob, which stores a big file are several smaller\npieces, might solve your problem. Unfortunately the topic is not\ncomplete yet.\n-- \nDuy\n"},{"id":"185326","messageId":"20120224101121.GA9526@pomac.netswarm.net","threadId":"29693","inReplyTo":"CACsJy8Bdbegs7QdztvsFnKPcpAX5UL7s7uc37wF3_nF4kJQjrQ@mail.gmail.com","subject":"Re: [FYI] very large text files and their problems.","fromName":"Ian Kumlien","fromEmail":"pomac@vapor.com","sentAt":"2012-02-24T10:11:21Z","receivedAt":"2012-02-24T10:11:21Z","isPatch":false,"sender":{"key":"pomac@vapor.com","avatar":null},"body":"I'm uncertain if you got my reply since i did it out of bounds - so i'll\nrepeat myself - sorry... =)\n\nOn Wed, Feb 22, 2012 at 11:18:19PM +0700, Nguyen Thai Ngoc Duy wrote:\n> On Wed, Feb 22, 2012 at 10:49 PM, Ian Kumlien <pomac@vapor.com> wrote:\n> > Hi,\n> >\n> > We just saw a interesting issue, git compressed a ~3.4 gb project to ~57 mb.\n> \n> How big are those files? How many of them? How often do they change?\n\nThis was the first check in, there is no deltas yet.\n\nThe file in question is ~3.3 gb in size - ie exactly: 3310214313 bytes\n(as seen below in the malloc failure)\n\ngit show <blob sha1 id> |wc -c gives the same exact result.\n\n> > But when we tried to clone it on a big machine we got:\n> >\n> > fatal: Out of memory, malloc failed (tried to allocate\n> > 18446744072724798634 bytes)\n> >\n> > This is already fixed in the 1.7.10 mainline - but it also seems like\n> \n> Does 1.7.9 have this problem?\n\nOnly tested 1.7.8 and 1.7.9.1 - works in mainline git (pre-1.7.10)\n\n> > git needs to have atleast the same ammount of memory as the largest\n> > file free... Couldn't this be worked around?\n> >\n> > On a (32 bit) machine with 4GB memory - results in:\n> > fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)\n> >\n> > (and i see how this could be a problem, but couldn't it be mitigated? or\n> > is it bydesign and intended behaviour?)\n> \n> I think that it's delta resolving that hogs all your memory. If your\n> files are smaller than 512M, try lower core.bigFileThreshold. The\n> topic jc/split-blob, which stores a big file are several smaller\n> pieces, might solve your problem. Unfortunately the topic is not\n> complete yet.\n\nWell, in this case it's just stream unpacking gzip data to disk, i\nunderstand if delta would be a problem... But wouldn't delta be a\nproblem in the sence of <size_of_change>+<size_of_subdata>+<result> ?\n\nIe, if the file is mmapped - it shouldn't have to be allocated, right?\n\n> -- \n> Duy\n"},{"id":"185329","messageId":"CACsJy8Aq=ofaHo3RkOBWTYP3eehZAU0E=HeDoUpR0nh+seKxzA@mail.gmail.com","threadId":"29693","inReplyTo":"20120224101121.GA9526@pomac.netswarm.net","subject":"Re: [FYI] very large text files and their problems.","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-02-24T11:14:46Z","receivedAt":"2012-02-24T11:14:46Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Feb 24, 2012 at 5:11 PM, Ian Kumlien <pomac@vapor.com> wrote:\n> I'm uncertain if you got my reply since i did it out of bounds - so i'll\n> repeat myself - sorry... =)\n\nyes I received it, just too busy this week.\n\n>> > git needs to have atleast the same ammount of memory as the largest\n>> > file free... Couldn't this be worked around?\n>> >\n>> > On a (32 bit) machine with 4GB memory - results in:\n>> > fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)\n>> >\n>> > (and i see how this could be a problem, but couldn't it be mitigated? or\n>> > is it bydesign and intended behaviour?)\n>>\n>> I think that it's delta resolving that hogs all your memory. If your\n>> files are smaller than 512M, try lower core.bigFileThreshold. The\n>> topic jc/split-blob, which stores a big file are several smaller\n>> pieces, might solve your problem. Unfortunately the topic is not\n>> complete yet.\n>\n> Well, in this case it's just stream unpacking gzip data to disk, i\n> understand if delta would be a problem... But wouldn't delta be a\n> problem in the sence of <size_of_change>+<size_of_subdata>+<result> ?\n>\n> Ie, if the file is mmapped - it shouldn't have to be allocated, right?\n\nWe should not delta large files. I was worried that the large file\ncheck could go wrong, But I guess your blob's not deltified in this\ncase.\n\nWhen you receive a pack during a clone, the pack is streamed to\nindex-pack, not mmapped, and index-pack checks every object in there\nin uncompressed form. I think I have found a way to avoid allocating\nthat much. Need some more check, then send out.\n-- \nDuy\n"},{"id":"185332","messageId":"20120224125552.GB9526@pomac.netswarm.net","threadId":"29693","inReplyTo":"CACsJy8Aq=ofaHo3RkOBWTYP3eehZAU0E=HeDoUpR0nh+seKxzA@mail.gmail.com","subject":"Re: [FYI] very large text files and their problems.","fromName":"Ian Kumlien","fromEmail":"pomac@vapor.com","sentAt":"2012-02-24T12:55:52Z","receivedAt":"2012-02-24T12:55:52Z","isPatch":false,"sender":{"key":"pomac@vapor.com","avatar":null},"body":"On Fri, Feb 24, 2012 at 06:14:46PM +0700, Nguyen Thai Ngoc Duy wrote:\n> On Fri, Feb 24, 2012 at 5:11 PM, Ian Kumlien <pomac@vapor.com> wrote:\n> > I'm uncertain if you got my reply since i did it out of bounds - so i'll\n> > repeat myself - sorry... =)\n> \n> yes I received it, just too busy this week.\n\nAh good, you never know what anti-spam measures people applies these\ndays... =)\n\nAnd i have the same, so i totally understand.\n\n> >> > git needs to have atleast the same ammount of memory as the largest\n> >> > file free... Couldn't this be worked around?\n> >> >\n> >> > On a (32 bit) machine with 4GB memory - results in:\n> >> > fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)\n> >> >\n> >> > (and i see how this could be a problem, but couldn't it be mitigated? or\n> >> > is it bydesign and intended behaviour?)\n> >>\n> >> I think that it's delta resolving that hogs all your memory. If your\n> >> files are smaller than 512M, try lower core.bigFileThreshold. The\n> >> topic jc/split-blob, which stores a big file are several smaller\n> >> pieces, might solve your problem. Unfortunately the topic is not\n> >> complete yet.\n> >\n> > Well, in this case it's just stream unpacking gzip data to disk, i\n> > understand if delta would be a problem... But wouldn't delta be a\n> > problem in the sence of <size_of_change>+<size_of_subdata>+<result> ?\n> >\n> > Ie, if the file is mmapped - it shouldn't have to be allocated, right?\n> \n> We should not delta large files. I was worried that the large file\n> check could go wrong, But I guess your blob's not deltified in this\n> case.\n\nThat would be correct\n\n> When you receive a pack during a clone, the pack is streamed to\n> index-pack, not mmapped, and index-pack checks every object in there\n> in uncompressed form. I think I have found a way to avoid allocating\n> that much. Need some more check, then send out.\n\nAh! That explains alot - do you have a publicly available version i\ncould look at?\n\n> -- \n> Duy\n"}]}