threads / discuss / 29693

[FYI] very large text files and their problems.

Subject: [FYI] very large text files and their problems.

## tl;dr

5 messages between Feb 22, 2012 and Feb 24, 2012.

replies: 4people: 2as markdown or json

Ian Kumlien· Feb 22, 2012, 15:49 UTC · lore
Hi, 

We just saw a interesting issue, git compressed a ~3.4 gb project to ~57 mb. But when we tried to clone it on a big machine we got:

fatal: Out of memory, malloc failed (tried to allocate 18446744072724798634 bytes)

This is already fixed in the 1.7.10 mainline - but it also seems like git needs to have atleast the same ammount of memory as the largest file free... Couldn't this be worked around?

On a (32 bit) machine with 4GB memory - results in: fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)

(and i see how this could be a problem, but couldn't it be mitigated? or is it bydesign and intended behaviour?)

I'm not subscribed to please keep me in CC.
/Ian Kumlien
Nguyen Thai Ngoc Duy· Feb 22, 2012, 16:18 UTC · re: Ian Kumlien · lore

Re: [FYI] very large text files and their problems.

On Wed, Feb 22, 2012 at 10:49 PM, Ian Kumlien <pomac@vapor.com> wrote:
> Hi,
>
> We just saw a interesting issue, git compressed a ~3.4 gb project to ~57 mb.
How big are those files? How many of them? How often do they change?
Show 6 quoted lines
> But when we tried to clone it on a big machine we got:
>
> fatal: Out of memory, malloc failed (tried to allocate
> 18446744072724798634 bytes)
>
> This is already fixed in the 1.7.10 mainline - but it also seems like
Does 1.7.9 have this problem?
Show 8 quoted lines
> git needs to have atleast the same ammount of memory as the largest
> file free... Couldn't this be worked around?
>
> On a (32 bit) machine with 4GB memory - results in:
> fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)
>
> (and i see how this could be a problem, but couldn't it be mitigated? or
> is it bydesign and intended behaviour?)

I think that it's delta resolving that hogs all your memory. If your files are smaller than 512M, try lower core.bigFileThreshold. The topic jc/split-blob, which stores a big file are several smaller pieces, might solve your problem. Unfortunately the topic is not complete yet.

-- 
Duy
Ian Kumlien· Feb 24, 2012, 10:11 UTC · re: Nguyen Thai Ngoc Duy · lore

Re: [FYI] very large text files and their problems.

I'm uncertain if you got my reply since i did it out of bounds - so i'll repeat myself - sorry... =)

On Wed, Feb 22, 2012 at 11:18:19PM +0700, Nguyen Thai Ngoc Duy wrote:
Show 6 quoted lines
> On Wed, Feb 22, 2012 at 10:49 PM, Ian Kumlien <pomac@vapor.com> wrote:
> > Hi,
> >
> > We just saw a interesting issue, git compressed a ~3.4 gb project to ~57 mb.
> 
> How big are those files? How many of them? How often do they change?
This was the first check in, there is no deltas yet.

The file in question is ~3.3 gb in size - ie exactly: 3310214313 bytes (as seen below in the malloc failure)

git show <blob sha1 id> |wc -c gives the same exact result.
Show 8 quoted lines
> > But when we tried to clone it on a big machine we got:
> >
> > fatal: Out of memory, malloc failed (tried to allocate
> > 18446744072724798634 bytes)
> >
> > This is already fixed in the 1.7.10 mainline - but it also seems like
> 
> Does 1.7.9 have this problem?
Only tested 1.7.8 and 1.7.9.1 - works in mainline git (pre-1.7.10)
Show 14 quoted lines
> > git needs to have atleast the same ammount of memory as the largest
> > file free... Couldn't this be worked around?
> >
> > On a (32 bit) machine with 4GB memory - results in:
> > fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)
> >
> > (and i see how this could be a problem, but couldn't it be mitigated? or
> > is it bydesign and intended behaviour?)
> 
> I think that it's delta resolving that hogs all your memory. If your
> files are smaller than 512M, try lower core.bigFileThreshold. The
> topic jc/split-blob, which stores a big file are several smaller
> pieces, might solve your problem. Unfortunately the topic is not
> complete yet.

Well, in this case it's just stream unpacking gzip data to disk, i understand if delta would be a problem... But wouldn't delta be a problem in the sence of <size_of_change>+<size_of_subdata>+<result> ?

Ie, if the file is mmapped - it shouldn't have to be allocated, right?
> -- 
> Duy
Nguyen Thai Ngoc Duy· Feb 24, 2012, 11:14 UTC · re: Ian Kumlien · lore

Re: [FYI] very large text files and their problems.

On Fri, Feb 24, 2012 at 5:11 PM, Ian Kumlien <pomac@vapor.com> wrote:
> I'm uncertain if you got my reply since i did it out of bounds - so i'll
> repeat myself - sorry... =)
yes I received it, just too busy this week.
Show 20 quoted lines
>> > git needs to have atleast the same ammount of memory as the largest
>> > file free... Couldn't this be worked around?
>> >
>> > On a (32 bit) machine with 4GB memory - results in:
>> > fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)
>> >
>> > (and i see how this could be a problem, but couldn't it be mitigated? or
>> > is it bydesign and intended behaviour?)
>>
>> I think that it's delta resolving that hogs all your memory. If your
>> files are smaller than 512M, try lower core.bigFileThreshold. The
>> topic jc/split-blob, which stores a big file are several smaller
>> pieces, might solve your problem. Unfortunately the topic is not
>> complete yet.
>
> Well, in this case it's just stream unpacking gzip data to disk, i
> understand if delta would be a problem... But wouldn't delta be a
> problem in the sence of <size_of_change>+<size_of_subdata>+<result> ?
>
> Ie, if the file is mmapped - it shouldn't have to be allocated, right?

We should not delta large files. I was worried that the large file check could go wrong, But I guess your blob's not deltified in this case.

When you receive a pack during a clone, the pack is streamed to index-pack, not mmapped, and index-pack checks every object in there in uncompressed form. I think I have found a way to avoid allocating that much. Need some more check, then send out.

-- 
Duy
Ian Kumlien· Feb 24, 2012, 12:55 UTC · re: Nguyen Thai Ngoc Duy · lore

Re: [FYI] very large text files and their problems.

On Fri, Feb 24, 2012 at 06:14:46PM +0700, Nguyen Thai Ngoc Duy wrote:
Show 5 quoted lines
> On Fri, Feb 24, 2012 at 5:11 PM, Ian Kumlien <pomac@vapor.com> wrote:
> > I'm uncertain if you got my reply since i did it out of bounds - so i'll
> > repeat myself - sorry... =)
> 
> yes I received it, just too busy this week.

Ah good, you never know what anti-spam measures people applies these days... =)

And i have the same, so i totally understand.
Show 24 quoted lines
> >> > git needs to have atleast the same ammount of memory as the largest
> >> > file free... Couldn't this be worked around?
> >> >
> >> > On a (32 bit) machine with 4GB memory - results in:
> >> > fatal: Out of memory, malloc failed (tried to allocate 3310214313 bytes)
> >> >
> >> > (and i see how this could be a problem, but couldn't it be mitigated? or
> >> > is it bydesign and intended behaviour?)
> >>
> >> I think that it's delta resolving that hogs all your memory. If your
> >> files are smaller than 512M, try lower core.bigFileThreshold. The
> >> topic jc/split-blob, which stores a big file are several smaller
> >> pieces, might solve your problem. Unfortunately the topic is not
> >> complete yet.
> >
> > Well, in this case it's just stream unpacking gzip data to disk, i
> > understand if delta would be a problem... But wouldn't delta be a
> > problem in the sence of <size_of_change>+<size_of_subdata>+<result> ?
> >
> > Ie, if the file is mmapped - it shouldn't have to be allocated, right?
> 
> We should not delta large files. I was worried that the large file
> check could go wrong, But I guess your blob's not deltified in this
> case.
That would be correct
> When you receive a pack during a clone, the pack is streamed to
> index-pack, not mmapped, and index-pack checks every object in there
> in uncompressed form. I think I have found a way to avoid allocating
> that much. Need some more check, then send out.

Ah! That explains alot - do you have a publicly available version i could look at?

> -- 
> Duy

← back to recent threads