{"thread":{"id":"35534","subject":"\"git fsck\" fails on malloc of 80 G","startedAt":"2013-12-16T16:05:32Z","lastAt":"2013-12-18T22:09:23Z","messageCount":6,"participants":["Dale R. Worley","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"232053","messageId":"201312161605.rBGG5Wm5027739@hobgoblin.ariadne.com","threadId":"35534","inReplyTo":null,"subject":"\"git fsck\" fails on malloc of 80 G","fromName":"Dale R. Worley","fromEmail":"worley@alum.mit.edu","sentAt":"2013-12-16T16:05:32Z","receivedAt":"2013-12-16T16:05:32Z","isPatch":false,"sender":{"key":"worley@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/19911107?v=4"},"body":"I have a large repository (17 GiB of disk used), although no single\nfile in the repository is over 1 GiB.  (I have pack.packSizeLimit set\nto \"1g\".)  I don't know how many files are in the repository, but it\nshouldn't exceed several tens of commits each containing several tens\nof thousands of files.\n\nDue to Git crashing while performing an operation, I want to verify\nthat the repository is consistent.  However, when I run \"git fsck\" it\nfails, apparently because it is trying to allocate 80 G of memory.  (I\ncan still do adds, commits, etc.)\n\n# git fsck\nChecking object directories: 100% (256/256), done.\nfatal: Out of memory, malloc failed (tried to allocate 80530636801 bytes)\n#\n\nI don't know if this is due to an outright bug or not.  But it seems\nto me that \"git fsck\" should not need to allocate any more memory than\nthe size (1 GiB) of a single pack file.  And given its purpose, \"git\nfsck\" should be one of the *most* robust Git tools!\n\nDale\n"},{"id":"232063","messageId":"20131216191500.GD29324@sigill.intra.peff.net","threadId":"35534","inReplyTo":"201312161605.rBGG5Wm5027739@hobgoblin.ariadne.com","subject":"Re: \"git fsck\" fails on malloc of 80 G","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-16T19:15:00Z","receivedAt":"2013-12-16T19:15:00Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Dec 16, 2013 at 11:05:32AM -0500, Dale R. Worley wrote:\n\n> # git fsck\n> Checking object directories: 100% (256/256), done.\n> fatal: Out of memory, malloc failed (tried to allocate 80530636801 bytes)\n> #\n\nCan you give you give us a backtrace from the die() call? It would help\nto know what it was trying to allocate 80G for.\n\n> I don't know if this is due to an outright bug or not.  But it seems\n> to me that \"git fsck\" should not need to allocate any more memory than\n> the size (1 GiB) of a single pack file.  And given its purpose, \"git\n> fsck\" should be one of the *most* robust Git tools!\n\nAgreed. Fsck tends to be more robust, but there are still many code\npaths that can die(). One of the problems I ran into recently is that\ncorrupt data can cause it to make a large allocation; we notice the\nbogus data as soon as we try to start filling the buffer, but sometimes\nthe bogus allocation is large enough to kill the process.\n\nThat was fixed by b039718, which is in master but not yet any released\nversion. You might see whether that helps.\n\n-Peff\n"},{"id":"232134","messageId":"201312180306.rBI36KCm016209@hobgoblin.ariadne.com","threadId":"35534","inReplyTo":"20131216191500.GD29324@sigill.intra.peff.net","subject":"Re: \"git fsck\" fails on malloc of 80 G","fromName":"Dale R. Worley","fromEmail":"worley@alum.mit.edu","sentAt":"2013-12-18T03:06:20Z","receivedAt":"2013-12-18T03:06:20Z","isPatch":false,"sender":{"key":"worley@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/19911107?v=4"},"body":"> From: Jeff King <peff@peff.net>\n> \n> On Mon, Dec 16, 2013 at 11:05:32AM -0500, Dale R. Worley wrote:\n> \n> > # git fsck\n> > Checking object directories: 100% (256/256), done.\n> > fatal: Out of memory, malloc failed (tried to allocate 80530636801 bytes)\n> > #\n> \n> Can you give you give us a backtrace from the die() call? It would help\n> to know what it was trying to allocate 80G for.\n\nFurther information:\n\n    # git --version\n    git version 1.8.3.1\n    #\n\nHere's the basic backtrace information, and the values of the \"size\"\nvariables, which seem to be the immediate culprits:\n\n    # gdb\n    GNU gdb (GDB) Fedora 7.6.1-46.fc19\n    Copyright (C) 2013 Free Software Foundation, Inc.\n    License GPLv3+: GNU GPL version 3 or later <http://gnu.org/licenses/gpl.html>\n    This is free software: you are free to change and redistribute it.\n    There is NO WARRANTY, to the extent permitted by law.  Type \"show copying\"\n    and \"show warranty\" for details.\n    This GDB was configured as \"x86_64-redhat-linux-gnu\".\n    For bug reporting instructions, please see:\n    <http://www.gnu.org/software/gdb/bugs/>.\n    (gdb) file /usr/bin/git\n    Reading symbols from /usr/bin/git...Reading symbols from /usr/lib/debug/usr/bin/git.debug...done.\n    done.\n    (gdb) break wrapper.c:59\n    Breakpoint 1 at 0x4f35ef: file wrapper.c, line 59.\n    (gdb) break die_child\n    Breakpoint 2 at 0x4d0ca0: file run-command.c, line 211.\n    (gdb) break die_async\n    Breakpoint 3 at 0x4d1020: file run-command.c, line 604.\n    (gdb) run fsck\n    Starting program: /usr/bin/git fsck\n    [Thread debugging using libthread_db enabled]\n    Using host libthread_db library \"/lib64/libthread_db.so.1\".\n    Checking object directories: 100% (256/256), done.\n    Checking objects:   0% (0/526211)   \n    Breakpoint 1, xmalloc (size=size@entry=80530636801) at wrapper.c:59\n    59\t\t\t\tdie(\"Out of memory, malloc failed (tried to allocate %lu bytes)\",\n    (gdb) bt\n    #0  xmalloc (size=size@entry=80530636801) at wrapper.c:59\n    #1  0x00000000004f3633 in xmallocz (size=size@entry=80530636800)\n\tat wrapper.c:73\n    #2  0x00000000004d922f in unpack_compressed_entry (p=p@entry=0x7e4020, \n\tw_curs=w_curs@entry=0x7fffffffc9f0, curpos=654214694, size=80530636800)\n\tat sha1_file.c:1797\n    #3  0x00000000004db4cb in unpack_entry (p=p@entry=0x7e4020, \n\tobj_offset=654214688, final_type=final_type@entry=0x7fffffffd088, \n\tfinal_size=final_size@entry=0x7fffffffd098) at sha1_file.c:2072\n    #4  0x00000000004b1e3f in verify_packfile (base_count=0, progress=0x9bdd80, \n\tfn=0x42fc00 <fsck_obj_buffer>, w_curs=0x7fffffffd090, p=0x7e4020)\n\tat pack-check.c:119\n    #5  verify_pack (p=p@entry=0x7e4020, fn=fn@entry=0x42fc00 <fsck_obj_buffer>, \n\tprogress=0x9bdd80, base_count=base_count@entry=0) at pack-check.c:177\n    #6  0x0000000000430724 in cmd_fsck (argc=0, argv=0x7fffffffe400, \n\tprefix=<optimized out>) at builtin/fsck.c:678\n    #7  0x0000000000405cfd in run_builtin (argv=0x7fffffffe400, argc=1, \n\tp=0x75fa68 <commands.23748+840>) at git.c:284\n    #8  handle_internal_command (argc=1, argv=0x7fffffffe400) at git.c:446\n    #9  0x000000000040511f in run_argv (argv=0x7fffffffe2a0, argcp=0x7fffffffe2ac)\n\tat git.c:492\n    #10 main (argc=1, argv=0x7fffffffe400) at git.c:567\n    (gdb) frame 2\n    #2  0x00000000004d922f in unpack_compressed_entry (p=p@entry=0x7e4020, \n\tw_curs=w_curs@entry=0x7fffffffc9f0, curpos=654214694, size=80530636800)\n\tat sha1_file.c:1797\n    1797\t\tbuffer = xmallocz(size);\n    (gdb) p size\n    $29 = 80530636800\n    (gdb) p/x size\n    $30 = 0x12c0000000\n    (gdb) frame 3\n    #3  0x00000000004db4cb in unpack_entry (p=p@entry=0x7e4020, \n\tobj_offset=654214688, final_type=final_type@entry=0x7fffffffd088, \n\tfinal_size=final_size@entry=0x7fffffffd098) at sha1_file.c:2072\n    2072\t\t\t\tdata = unpack_compressed_entry(p, &w_curs, curpos, size);\n    (gdb) p size\n    $31 = 80530636800\n    (gdb) p/x size\n    $32 = 0x12c0000000\n    (gdb) \n\nI did a further test to see where the value of \"size\" came from:\n\n    (gdb) break sha1_file.c:2023\n    Breakpoint 4 at 0x4db073: file sha1_file.c, line 2023.\n    (gdb) cond 4 size == 0x12c0000000\n    (gdb) break sha1_file.c:2029\n    Breakpoint 5 at 0x4daee7: file sha1_file.c, line 2029.\n    (gdb) cond 5 size == 0x12c0000000\n    (gdb) break sha1_file.c:2072\n    Breakpoint 6 at 0x4db4b4: file sha1_file.c, line 2072.\n    (gdb) cond 6 size == 0x12c0000000\n    (gdb) break unpack_object_header_buffer\n    Breakpoint 7 at 0x4d9ea0: file sha1_file.c, line 1399.\n    (gdb) comm 7\n    Type commands for breakpoint(s) 7, one per line.\n    End with a line saying just \"end\".\n    >continue\n    >end\n    (gdb) run\n    The program being debugged has been started already.\n    Start it from the beginning? (y or n) y\n    Starting program: /usr/bin/git fsck\n    [Thread debugging using libthread_db enabled]\n    Using host libthread_db library \"/lib64/libthread_db.so.1\".\n    Checking object directories: 100% (256/256), done.\n\n    Breakpoint 7, unpack_object_header_buffer (\n\tbuf=0x7fffc4d3e00c \"\\265\\334\\352\\277\\023x\\234\", len=733530087, \n\ttype=type@entry=0x7fffffffc984, sizep=sizep@entry=0x7fffffffca00)\n\tat sha1_file.c:1399\n    1399\t{\n    Checking objects:   0% (0/526211)   \n    Breakpoint 7, unpack_object_header_buffer (\n\tbuf=0x7fffebd26620 \"\\260\\200\\200\\200\\340\\022x\\234\\354\\301\\001\\001\", \n\tlen=79315411, type=type@entry=0x7fffffffc984, \n\tsizep=sizep@entry=0x7fffffffca00) at sha1_file.c:1399\n    1399\t{\n\n    Breakpoint 5, unpack_entry (p=p@entry=0x7e4020, obj_offset=654214688, \n\tfinal_type=final_type@entry=0x7fffffffd088, \n\tfinal_size=final_size@entry=0x7fffffffd098) at sha1_file.c:2029\n    2029\t\t\tif (type != OBJ_OFS_DELTA && type != OBJ_REF_DELTA)\n    (gdb) \n\nIf I understand the code correctly, the object header buffer\n\\260\\200\\200\\200\\340\\022x\\234\\354\\301\\001\\001\nreally does encode the size value 0x12c0000000.\n\nI will see if I can experiment with the new version you mention.\n\nDale\n"},{"id":"232210","messageId":"201312182108.rBIL8lAo015570@hobgoblin.ariadne.com","threadId":"35534","inReplyTo":"20131216191500.GD29324@sigill.intra.peff.net","subject":"Re: \"git fsck\" fails on malloc of 80 G","fromName":"Dale R. Worley","fromEmail":"worley@alum.mit.edu","sentAt":"2013-12-18T21:08:47Z","receivedAt":"2013-12-18T21:08:47Z","isPatch":false,"sender":{"key":"worley@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/19911107?v=4"},"body":"> From: Jeff King <peff@peff.net>\n\n> One of the problems I ran into recently is that\n> corrupt data can cause it to make a large allocation\n\nOne thing I notice is that in unpack_compressed_entry() in\nsha1_file.c, there is a mallocz of \"size\" bytes.  It appears that\n\"size\" is the size of the object that is being unpacked.  If so, this\ncode cannot be correct, because it assumes that any file that is\nstored in the repository can be put into a buffer allocated in RAM.\n\nDale\n"},{"id":"232208","messageId":"20131218215821.GA14276@sigill.intra.peff.net","threadId":"35534","inReplyTo":"201312180306.rBI36KCm016209@hobgoblin.ariadne.com","subject":"Re: \"git fsck\" fails on malloc of 80 G","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-18T21:58:21Z","receivedAt":"2013-12-18T21:58:21Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Dec 17, 2013 at 10:06:20PM -0500, Dale R. Worley wrote:\n\n> Here's the basic backtrace information, and the values of the \"size\"\n> variables, which seem to be the immediate culprits:\n> [...]\n>     #1  0x00000000004f3633 in xmallocz (size=size@entry=80530636800)\n> \tat wrapper.c:73\n>     #2  0x00000000004d922f in unpack_compressed_entry (p=p@entry=0x7e4020, \n> \tw_curs=w_curs@entry=0x7fffffffc9f0, curpos=654214694, size=80530636800)\n> \tat sha1_file.c:1797\n>     #3  0x00000000004db4cb in unpack_entry (p=p@entry=0x7e4020, \n> \tobj_offset=654214688, final_type=final_type@entry=0x7fffffffd088, \n> \tfinal_size=final_size@entry=0x7fffffffd098) at sha1_file.c:2072\n>     #4  0x00000000004b1e3f in verify_packfile (base_count=0, progress=0x9bdd80, \n> \tfn=0x42fc00 <fsck_obj_buffer>, w_curs=0x7fffffffd090, p=0x7e4020)\n> \tat pack-check.c:119\n\nThanks, that's helpful. Unfortunately the patch I mentioned before won't\nhelp you. The packfile format (like the experimental loose format that my patch\ndropped) stores the size outside of the zlib crc. So it has the same\nproblem: we want to allocate the buffer up front to store the zlib\nresults.\n\nThe pack index does store a crc (calculated when we made or received\nthe pack) over each object's on-disk representation. So we could check\nthat, though doing it on every access has performance implications.\n\nThe pack data itself also has a SHA-1 checksum over the whole thing. We\nshould probably do a better job in verify-pack of:\n\n  1. Check the whole sha1 checksum before doing anything else.\n\n  2. In the uncommon case that it fails, check each individual object\n     crc to find the broken object (and if none, assume either the\n     header or the checksum itself is what got munged).\n\nIn the meantime, you should be able to do step 1 manually like:\n\n  # check first N-20 bytes of packfile against the checksum in the\n  # final 20 bytes. NB: pretty sure this use of \"head\" is a GNU-ism,\n  # and of course you need openssl\n  for i in objects/pack/*.pack; do\n    tail -c 20 \"$i\" >want.tmp &&\n    head -c -20 \"$i\" | openssl sha1 -binary >have.tmp &&\n    cmp want.tmp have.tmp ||\n    echo >&2 \"broken: $i\"\n  done\n\ngit-fsck should be doing this check itself, but I wonder if you are not\nmaking it that far.\n\n> If I understand the code correctly, the object header buffer\n> \\260\\200\\200\\200\\340\\022x\\234\\354\\301\\001\\001\n> really does encode the size value 0x12c0000000.\n\nIf it does, and you do not have an 80G file, then it sounds like you may\nhave a corrupt packfile.\n\n-Peff\n"},{"id":"232209","messageId":"20131218220922.GA16347@sigill.intra.peff.net","threadId":"35534","inReplyTo":"201312182108.rBIL8lAo015570@hobgoblin.ariadne.com","subject":"Re: \"git fsck\" fails on malloc of 80 G","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-18T22:09:23Z","receivedAt":"2013-12-18T22:09:23Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Dec 18, 2013 at 04:08:47PM -0500, Dale R. Worley wrote:\n\n> > From: Jeff King <peff@peff.net>\n> \n> > One of the problems I ran into recently is that\n> > corrupt data can cause it to make a large allocation\n> \n> One thing I notice is that in unpack_compressed_entry() in\n> sha1_file.c, there is a mallocz of \"size\" bytes.  It appears that\n> \"size\" is the size of the object that is being unpacked.  If so, this\n> code cannot be correct, because it assumes that any file that is\n> stored in the repository can be put into a buffer allocated in RAM.\n\nFor some definition of correct. Git does load whole-blobs into memory in\nseveral places. Some code paths _can_ stream, but they do not stream\ndeltas, and the diff engine definitely wants the whole thing in-core.\n\nSo you are reading it right. If you want to work on changing it, be my\nguest, but it's a non-trivial fix. ;)\n\n-Peff\n"}]}