{"thread":{"id":"17392","subject":"\"malloc failed\"","startedAt":"2009-01-27T15:04:42Z","lastAt":"2009-01-30T04:49:19Z","messageCount":15,"participants":["David Abrahams","Shawn O. Pearce","Johannes Schindelin","Jeff King","Pau Garcia i Quiles","Junio C Hamano","Andreas Ericsson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"102122","messageId":"878wow7pth.fsf@mcbain.luannocracy.com","threadId":"17392","inReplyTo":null,"subject":"\"malloc failed\"","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-01-27T15:04:42Z","receivedAt":"2009-01-27T15:04:42Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nI've been abusing Git for a purpose it wasn't intended to serve:\narchiving a large number of files with many duplicates and\nnear-duplicates.  Every once in a while, when trying to do something\nreally big, it tells me \"malloc failed\" and bails out (I think it's\nduring \"git add\" but because of the way I issued the commands I can't\ntell: it could have been a commit or a gc).  This is on a 64-bit linux\nmachine with 8G of ram and plenty of swap space, so I'm surprised.\n\nGit is doing an amazing job at archiving and compressing all this stuff\nI'm putting in it, but I have to do it a wee bit at a time or it craps\nout.  Bug?\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"102125","messageId":"20090127152915.GA1321@spearce.org","threadId":"17392","inReplyTo":"878wow7pth.fsf@mcbain.luannocracy.com","subject":"Re: \"malloc failed\"","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-01-27T15:29:16Z","receivedAt":"2009-01-27T15:29:16Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"David Abrahams <dave@boostpro.com> wrote:\n> I've been abusing Git for a purpose it wasn't intended to serve:\n> archiving a large number of files with many duplicates and\n> near-duplicates.  Every once in a while, when trying to do something\n> really big, it tells me \"malloc failed\" and bails out (I think it's\n> during \"git add\" but because of the way I issued the commands I can't\n> tell: it could have been a commit or a gc).  This is on a 64-bit linux\n> machine with 8G of ram and plenty of swap space, so I'm surprised.\n> \n> Git is doing an amazing job at archiving and compressing all this stuff\n> I'm putting in it, but I have to do it a wee bit at a time or it craps\n> out.  Bug?\n\nNo, not really.  Above you said you are \"abusing git for a purpose\nit wasn't intended to serve\"...\n\nGit was never designed to handle many large binary blobs of data.\nIt was mostly designed for source code, where the majority of the\ndata stored in it is some form of text file written by a human.\n\nBy their very nature these files need to be relatively short (e.g.\nunder 1 MB each) as no human can sanely maintain a text file that\nlarge without breaking it apart into different smaller files (like\nthe source code for an operating system kernel).\n\nAs a result of this approach, the git code assumes it can malloc()\nat least two blocks large enough for each file: one of the fully\ndecompressed content, and another for the fully compressed content.\nTry doing git add on a large file and its very likely malloc\nwill fail due to ulimit issues, or you just don't have enough\nmemory/address space to go around.\n\ngit gc likewise needs a good chunk of memory, but it shouldn't\nusually report \"malloc failed\".  Usually in git gc if a malloc fails\nit prints a warning and degrades the quality of its data compression.\nBut there are critical bookkeeping data structures where we must be\nable to malloc the memory, and if those fail because we've already\nexhausted the heap early on, then yea, it can fail too.\n\n-- \nShawn.\n"},{"id":"102128","messageId":"87hc3k69y9.fsf@mcbain.luannocracy.com","threadId":"17392","inReplyTo":"20090127152915.GA1321@spearce.org","subject":"Re: \"malloc failed\"","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-01-27T15:32:46Z","receivedAt":"2009-01-27T15:32:46Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\non Tue Jan 27 2009, \"Shawn O. Pearce\" <spearce-AT-spearce.org> wrote:\n\n> David Abrahams <dave@boostpro.com> wrote:\n>> I've been abusing Git for a purpose it wasn't intended to serve:\n>> archiving a large number of files with many duplicates and\n>> near-duplicates.  Every once in a while, when trying to do something\n>> really big, it tells me \"malloc failed\" and bails out (I think it's\n>> during \"git add\" but because of the way I issued the commands I can't\n>> tell: it could have been a commit or a gc).  This is on a 64-bit linux\n>> machine with 8G of ram and plenty of swap space, so I'm surprised.\n>> \n>> Git is doing an amazing job at archiving and compressing all this stuff\n>> I'm putting in it, but I have to do it a wee bit at a time or it craps\n>> out.  Bug?\n>\n> No, not really.  Above you said you are \"abusing git for a purpose\n> it wasn't intended to serve\"...\n\nAbsolutely; I want to be upfront about that :-)\n\n> Git was never designed to handle many large binary blobs of data.\n\nThey're largely text blobs, although there definitely are a fair share\nof binaries.\n\n> It was mostly designed for source code, where the majority of the\n> data stored in it is some form of text file written by a human.\n>\n> By their very nature these files need to be relatively short (e.g.\n> under 1 MB each) as no human can sanely maintain a text file that\n> large without breaking it apart into different smaller files (like\n> the source code for an operating system kernel).\n>\n> As a result of this approach, the git code assumes it can malloc()\n> at least two blocks large enough for each file: one of the fully\n> decompressed content, and another for the fully compressed content.\n> Try doing git add on a large file and its very likely malloc\n> will fail due to ulimit issues, or you just don't have enough\n> memory/address space to go around.\n\nOh, so maybe I'm getting hit by ulimit; I didn't think of that.  I could\nraise my ulimit to try to get around this.\n\n> git gc likewise needs a good chunk of memory, but it shouldn't\n> usually report \"malloc failed\".  Usually in git gc if a malloc fails\n> it prints a warning and degrades the quality of its data compression.\n> But there are critical bookkeeping data structures where we must be\n> able to malloc the memory, and if those fail because we've already\n> exhausted the heap early on, then yea, it can fail too.\n\nThanks much for that, and for reminding me about ulimit.\n\nCheers,\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"102147","messageId":"alpine.DEB.1.00.0901271900260.3586@pacific.mpi-cbg.de","threadId":"17392","inReplyTo":"878wow7pth.fsf@mcbain.luannocracy.com","subject":"Re: \"malloc failed\"","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-01-27T18:02:42Z","receivedAt":"2009-01-27T18:02:42Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 27 Jan 2009, David Abrahams wrote:\n\n> I've been abusing Git for a purpose it wasn't intended to serve: \n> archiving a large number of files with many duplicates and \n> near-duplicates.\n\nHah!  My first UGFWIINI contender!  Unfortunately, I listed that purpose \nexplicitely already...\n\n> Every once in a while, when trying to do something really big, it tells \n> me \"malloc failed\" and bails out (I think it's during \"git add\" but \n> because of the way I issued the commands I can't tell: it could have \n> been a commit or a gc).  This is on a 64-bit linux machine with 8G of \n> ram and plenty of swap space, so I'm surprised.\n\nYes, I am surprised, too.  I would expect that some kind of arbitrary \nuser-specifiable limit hit you.  Haven't had time to look at the code, \nthough.\n\nCiao,\nDscho\n"},{"id":"102251","messageId":"20090128050225.GA18546@coredump.intra.peff.net","threadId":"17392","inReplyTo":"878wow7pth.fsf@mcbain.luannocracy.com","subject":"Re: \"malloc failed\"","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-01-28T05:02:25Z","receivedAt":"2009-01-28T05:02:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jan 27, 2009 at 10:04:42AM -0500, David Abrahams wrote:\n\n> I've been abusing Git for a purpose it wasn't intended to serve:\n> archiving a large number of files with many duplicates and\n> near-duplicates.  Every once in a while, when trying to do something\n> really big, it tells me \"malloc failed\" and bails out (I think it's\n> during \"git add\" but because of the way I issued the commands I can't\n> tell: it could have been a commit or a gc).  This is on a 64-bit linux\n> machine with 8G of ram and plenty of swap space, so I'm surprised.\n> \n> Git is doing an amazing job at archiving and compressing all this stuff\n> I'm putting in it, but I have to do it a wee bit at a time or it craps\n> out.  Bug?\n\nHow big is the repository? How big are the biggest files? I have a\n3.5G repo with files ranging from a few bytes to about 180M. I've never\nrun into malloc problems or gone into swap on my measly 1G box.\nHow does your dataset compare?\n\nAs others have mentioned, git wasn't really designed specifically for\nthose sorts of numbers, but in the interests of performance, I find git\nis usually pretty careful about not keeping too much useless stuff in\nmemory at one time.  And the fact that you can perform the same\noperation a little bit at a time and achieve success implies to me there\nmight be a leak or some silly behavior that can be fixed.\n\nIt would help a lot if we knew the operation that was causing the\nproblem. Can you try to isolate the failed command next time it happens?\n\n-Peff\n"},{"id":"102359","messageId":"c26bbb3fe074f6f6e0634a4ae8611239@206.71.190.141","threadId":"17392","inReplyTo":"20090128050225.GA18546@coredump.intra.peff.net","subject":"Re: \"malloc failed\"","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-01-28T21:53:49Z","receivedAt":"2009-01-28T21:53:49Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\nOn Wed, 28 Jan 2009 00:02:25 -0500, Jeff King <peff@peff.net> wrote:\n\n> On Tue, Jan 27, 2009 at 10:04:42AM -0500, David Abrahams wrote:\n\n> \n\n>> I've been abusing Git for a purpose it wasn't intended to serve:\n\n>> archiving a large number of files with many duplicates and\n\n>> near-duplicates.  Every once in a while, when trying to do something\n\n>> really big, it tells me \"malloc failed\" and bails out (I think it's\n\n>> during \"git add\" but because of the way I issued the commands I can't\n\n>> tell: it could have been a commit or a gc).  This is on a 64-bit linux\n\n>> machine with 8G of ram and plenty of swap space, so I'm surprised.\n\n>> \n\n>> Git is doing an amazing job at archiving and compressing all this stuff\n\n>> I'm putting in it, but I have to do it a wee bit at a time or it craps\n\n>> out.  Bug?\n\n> \n\n> How big is the repository? How big are the biggest files? I have a\n\n> 3.5G repo with files ranging from a few bytes to about 180M. I've never\n\n> run into malloc problems or gone into swap on my measly 1G box.\n\n> How does your dataset compare?\n\n\n\nI'll try to do some research.  Gotta go pick up my boy now...\n\n\n\n> As others have mentioned, git wasn't really designed specifically for\n\n> those sorts of numbers, but in the interests of performance, I find git\n\n> is usually pretty careful about not keeping too much useless stuff in\n\n> memory at one time.  And the fact that you can perform the same\n\n> operation a little bit at a time and achieve success implies to me there\n\n> might be a leak or some silly behavior that can be fixed.\n\n> \n\n> It would help a lot if we knew the operation that was causing the\n\n> problem. Can you try to isolate the failed command next time it happens?\n\n\n\nroot@recovery:/olympic/deuce/review# ulimit -v\n\nunlimited\n\nroot@recovery:/olympic/deuce/review# git add hydra.bak/home-dave\n\nfatal: Out of memory, malloc failed\n\n\n\n\n\nThe process never even gets close to my total installed RAM size, much less\n\nmy whole VM space size.\n\n\n\n-- \n\nDavid Abrahams\n\nBoostpro Computing\n\nhttp://www.boostpro.com\n"},{"id":"102361","messageId":"3af572ac0901281416x5adef0eak89bd4b40fda52c2b@mail.gmail.com","threadId":"17392","inReplyTo":"20090128050225.GA18546@coredump.intra.peff.net","subject":"Re: \"malloc failed\"","fromName":"Pau Garcia i Quiles","fromEmail":"pgquiles@elpauer.org","sentAt":"2009-01-28T22:16:32Z","receivedAt":"2009-01-28T22:16:32Z","isPatch":false,"sender":{"key":"pgquiles@elpauer.org","avatar":null},"body":"On Wed, Jan 28, 2009 at 6:02 AM, Jeff King <peff@peff.net> wrote:\n\n> How big is the repository? How big are the biggest files? I have a\n> 3.5G repo with files ranging from a few bytes to about 180M. I've never\n> run into malloc problems or gone into swap on my measly 1G box.\n> How does your dataset compare?\n\nI also have malloc problems but only on Windows, on Linux it works fine.\n\nMy case: I have a 500 MB repository with a 1GB working tree, with\nbinary files ranging from 100KB to 50MB and a few thousand source\nfiles.\n\nI have two branches ('master' and 'cmake') and the latter has suffered\na huge hierarchy reorganization.\n\nWhen I merge 'master' in 'cmake', if I use the 'subtree' strategy, it\nworks fine. If I use any other strategy, after a couple of minutes I\nreceive a \"malloc failed\" and the tree is all messed up. As I said, on\nLinux it works fine, so maybe it's a Windows-specific problem.\n\n-- \nPau Garcia i Quiles\nhttp://www.elpauer.org\n(Due to my workload, I may need 10 days to answer)\n"},{"id":"102374","messageId":"87skn3rn5n.fsf@mcbain.luannocracy.com","threadId":"17392","inReplyTo":"c26bbb3fe074f6f6e0634a4ae8611239@206.71.190.141","subject":"Re: \"malloc failed\"","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-01-29T00:06:28Z","receivedAt":"2009-01-29T00:06:28Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\non Wed Jan 28 2009, David Abrahams <dave-AT-boostpro.com> wrote:\n\n> On Wed, 28 Jan 2009 00:02:25 -0500, Jeff King <peff@peff.net> wrote:\n>> On Tue, Jan 27, 2009 at 10:04:42AM -0500, David Abrahams wrote:\n>> \n>>> I've been abusing Git for a purpose it wasn't intended to serve:\n>>> archiving a large number of files with many duplicates and\n>>> near-duplicates.  Every once in a while, when trying to do something\n>>> really big, it tells me \"malloc failed\" and bails out (I think it's\n>>> during \"git add\" but because of the way I issued the commands I can't\n>>> tell: it could have been a commit or a gc).  This is on a 64-bit linux\n>>> machine with 8G of ram and plenty of swap space, so I'm surprised.\n>>> \n>>> Git is doing an amazing job at archiving and compressing all this stuff\n>>> I'm putting in it, but I have to do it a wee bit at a time or it craps\n>>> out.  Bug?\n>> \n>> How big is the repository? How big are the biggest files? I have a\n>> 3.5G repo with files ranging from a few bytes to about 180M. I've never\n>> run into malloc problems or gone into swap on my measly 1G box.\n>> How does your dataset compare?\n>\n> I'll try to do some research.  Gotta go pick up my boy now...\n\nWell, moving the 2.6G .dar backup binary out of the fileset seems to\nhave helped a little, not surprisingly :-P\n\nI don't know whether anyone on this list should care about that failure\ngiven the level of abuse I'm inflicting on Git, but keep in mind that\nthe system *does* have 8G of memory.  Conclude what you will from that,\nI suppose!\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"102406","messageId":"20090129051451.GA31507@coredump.intra.peff.net","threadId":"17392","inReplyTo":"3af572ac0901281416x5adef0eak89bd4b40fda52c2b@mail.gmail.com","subject":"Re: \"malloc failed\"","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-01-29T05:14:51Z","receivedAt":"2009-01-29T05:14:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 28, 2009 at 11:16:32PM +0100, Pau Garcia i Quiles wrote:\n\n> My case: I have a 500 MB repository with a 1GB working tree, with\n> binary files ranging from 100KB to 50MB and a few thousand source\n> files.\n> \n> I have two branches ('master' and 'cmake') and the latter has suffered\n> a huge hierarchy reorganization.\n> \n> When I merge 'master' in 'cmake', if I use the 'subtree' strategy, it\n> works fine. If I use any other strategy, after a couple of minutes I\n> receive a \"malloc failed\" and the tree is all messed up. As I said, on\n> Linux it works fine, so maybe it's a Windows-specific problem.\n\nHmm. It very well might be the rename detection allocating a lot of\nmemory to do inexact rename detection. It does try to limit the amount\nof work, but based on number of files. So if you have a lot of huge\nfiles, that might be fooling it.\n\nTry setting merge.renamelimit to something small (but not '0', which\nmeans \"no limit\").\n\n-Peff\n"},{"id":"102407","messageId":"20090129052041.GB31507@coredump.intra.peff.net","threadId":"17392","inReplyTo":"87skn3rn5n.fsf@mcbain.luannocracy.com","subject":"Re: \"malloc failed\"","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-01-29T05:20:41Z","receivedAt":"2009-01-29T05:20:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jan 28, 2009 at 07:06:28PM -0500, David Abrahams wrote:\n\n> Well, moving the 2.6G .dar backup binary out of the fileset seems to\n> have helped a little, not surprisingly :-P\n\nOk, that _is_ big. ;) I wouldn't be surprised if there is some corner of\nthe code that barfs on a single object that doesn't fit in a signed\n32-bit integer; I don't think we have any test coverage for stuff that\nbig.\n\nBut it may also just be that we are going to try malloc'ing 2.6G, and\nthat's making some system limit unhappy.\n\n> I don't know whether anyone on this list should care about that failure\n> given the level of abuse I'm inflicting on Git, but keep in mind that\n> the system *does* have 8G of memory.  Conclude what you will from that,\n> I suppose!\n\nWell, I think you said before that you were never getting close to using\nup all your memory. Which implies it's some system limit.\n\n-Peff\n"},{"id":"102408","messageId":"20090129055633.GA32609@coredump.intra.peff.net","threadId":"17392","inReplyTo":"20090129052041.GB31507@coredump.intra.peff.net","subject":"Re: \"malloc failed\"","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-01-29T05:56:34Z","receivedAt":"2009-01-29T05:56:34Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 29, 2009 at 12:20:41AM -0500, Jeff King wrote:\n\n> Ok, that _is_ big. ;) I wouldn't be surprised if there is some corner of\n> the code that barfs on a single object that doesn't fit in a signed\n> 32-bit integer; I don't think we have any test coverage for stuff that\n> big.\n\nSure enough, that is the problem. With the patch below I was able to\n\"git add\" and commit a 3 gigabyte file of random bytes (so even the\ndeflated object was 3G).\n\nI think it might be worth applying as a general cleanup, but I have no\nidea if other parts of the system might barf on such an object.\n\n-- >8 --\nSubject: [PATCH] avoid 31-bit truncation in write_loose_object\n\nThe size of the content we are adding may be larger than\n2.1G (i.e., \"git add gigantic-file\"). Most of the code-path\nto do so uses size_t or unsigned long to record the size,\nbut write_loose_object uses a signed int.\n\nOn platforms where \"int\" is 32-bits (which includes x86_64\nLinux platforms), we end up passing malloc a negative size.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1_file.c |    3 ++-\n 1 files changed, 2 insertions(+), 1 deletions(-)\n\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 360f7e5..8868b80 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -2340,7 +2340,8 @@ static int create_tmpfile(char *buffer, size_t bufsiz, const char *filename)\n static int write_loose_object(const unsigned char *sha1, char *hdr, int hdrlen,\n \t\t\t      void *buf, unsigned long len, time_t mtime)\n {\n-\tint fd, size, ret;\n+\tint fd, ret;\n+\tsize_t size;\n \tunsigned char *compressed;\n \tz_stream stream;\n \tchar *filename;\n-- \n1.6.1.1.259.g8712.dirty\n"},{"id":"102414","messageId":"7vfxj2h7ka.fsf@gitster.siamese.dyndns.org","threadId":"17392","inReplyTo":"20090129055633.GA32609@coredump.intra.peff.net","subject":"Re: \"malloc failed\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-01-29T07:53:25Z","receivedAt":"2009-01-29T07:53:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Subject: [PATCH] avoid 31-bit truncation in write_loose_object\n>\n> The size of the content we are adding may be larger than\n> 2.1G (i.e., \"git add gigantic-file\"). Most of the code-path\n> to do so uses size_t or unsigned long to record the size,\n> but write_loose_object uses a signed int.\n\nThanks.\n\nI wonder if some analysis tool like sparse can help us spot these...\n"},{"id":"102443","messageId":"87pri6qmvm.fsf@mcbain.luannocracy.com","threadId":"17392","inReplyTo":"20090129055633.GA32609@coredump.intra.peff.net","subject":"Re: \"malloc failed\"","fromName":"David Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2009-01-29T13:10:05Z","receivedAt":"2009-01-29T13:10:05Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"\non Thu Jan 29 2009, Jeff King <peff-AT-peff.net> wrote:\n\n> On Thu, Jan 29, 2009 at 12:20:41AM -0500, Jeff King wrote:\n>\n>> Ok, that _is_ big. ;) I wouldn't be surprised if there is some corner of\n>> the code that barfs on a single object that doesn't fit in a signed\n>> 32-bit integer; I don't think we have any test coverage for stuff that\n>> big.\n>\n> Sure enough, that is the problem. With the patch below I was able to\n> \"git add\" and commit a 3 gigabyte file of random bytes (so even the\n> deflated object was 3G).\n>\n> I think it might be worth applying as a general cleanup, but I have no\n> idea if other parts of the system might barf on such an object.\n>\n> -- >8 --\n> Subject: [PATCH] avoid 31-bit truncation in write_loose_object\n>\n> The size of the content we are adding may be larger than\n> 2.1G (i.e., \"git add gigantic-file\"). Most of the code-path\n> to do so uses size_t or unsigned long to record the size,\n> but write_loose_object uses a signed int.\n>\n> On platforms where \"int\" is 32-bits (which includes x86_64\n> Linux platforms), we end up passing malloc a negative size.\n\n\nGood work.  I don't know if this matters to you, but I think on a 32-bit\nplatform you'll find that size_t, which is supposed to be able to hold\nthe size of the largest representable *memory block*, is only 4 bytes\nlarge:\n\n  #include <limits.h>\n  #include <stdio.h>\n\n  int main()\n  {\n    printf(\"sizeof(size_t) = %d\", sizeof(size_t));\n  }\n\nPrints \"sizeof(size_t) = 4\" on my core duo.\n\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n>  sha1_file.c |    3 ++-\n>  1 files changed, 2 insertions(+), 1 deletions(-)\n>\n> diff --git a/sha1_file.c b/sha1_file.c\n> index 360f7e5..8868b80 100644\n> --- a/sha1_file.c\n> +++ b/sha1_file.c\n> @@ -2340,7 +2340,8 @@ static int create_tmpfile(char *buffer, size_t bufsiz, const\n> char *filename)\n>  static int write_loose_object(const unsigned char *sha1, char *hdr, int hdrlen,\n>  \t\t\t      void *buf, unsigned long len, time_t mtime)\n>  {\n> -\tint fd, size, ret;\n> +\tint fd, ret;\n> +\tsize_t size;\n>  \tunsigned char *compressed;\n>  \tz_stream stream;\n>  \tchar *filename;\n\n-- \nDave Abrahams\nBoostPro Computing\nhttp://www.boostpro.com\n"},{"id":"102445","messageId":"4981B1FB.6030700@op5.se","threadId":"17392","inReplyTo":"87pri6qmvm.fsf@mcbain.luannocracy.com","subject":"Re: \"malloc failed\"","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-01-29T13:41:15Z","receivedAt":"2009-01-29T13:41:15Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"David Abrahams wrote:\n> on Thu Jan 29 2009, Jeff King <peff-AT-peff.net> wrote:\n> \n>> On Thu, Jan 29, 2009 at 12:20:41AM -0500, Jeff King wrote:\n>>\n>>> Ok, that _is_ big. ;) I wouldn't be surprised if there is some corner of\n>>> the code that barfs on a single object that doesn't fit in a signed\n>>> 32-bit integer; I don't think we have any test coverage for stuff that\n>>> big.\n>> Sure enough, that is the problem. With the patch below I was able to\n>> \"git add\" and commit a 3 gigabyte file of random bytes (so even the\n>> deflated object was 3G).\n>>\n>> I think it might be worth applying as a general cleanup, but I have no\n>> idea if other parts of the system might barf on such an object.\n>>\n>> -- >8 --\n>> Subject: [PATCH] avoid 31-bit truncation in write_loose_object\n>>\n>> The size of the content we are adding may be larger than\n>> 2.1G (i.e., \"git add gigantic-file\"). Most of the code-path\n>> to do so uses size_t or unsigned long to record the size,\n>> but write_loose_object uses a signed int.\n>>\n>> On platforms where \"int\" is 32-bits (which includes x86_64\n>> Linux platforms), we end up passing malloc a negative size.\n> \n> \n> Good work.  I don't know if this matters to you, but I think on a 32-bit\n> platform you'll find that size_t, which is supposed to be able to hold\n> the size of the largest representable *memory block*, is only 4 bytes\n> large:\n> \n>   #include <limits.h>\n>   #include <stdio.h>\n> \n>   int main()\n>   {\n>     printf(\"sizeof(size_t) = %d\", sizeof(size_t));\n>   }\n> \n> Prints \"sizeof(size_t) = 4\" on my core duo.\n> \n\nIt has nothing to do with typesize, and everything to do with\nsignedness. A size_t cannot be negative, while an int can.\nMaking sure we use the correct signedness everywhere means\nwe double the capacity where negative values are clearly bogus,\nsuch as in this case. On 32-bit platforms, the upper limit for\nwhat git can handle is now 4GB, which is expected. To go beyond\nthat, we'd need to rework the algorithm so we handle chunks of\nthe data instead of the whole. Some day, that might turn out to\nbe necessary but today is not that day.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"102542","messageId":"20090130044919.GA18655@coredump.intra.peff.net","threadId":"17392","inReplyTo":"87pri6qmvm.fsf@mcbain.luannocracy.com","subject":"Re: \"malloc failed\"","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-01-30T04:49:19Z","receivedAt":"2009-01-30T04:49:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 29, 2009 at 08:10:05AM -0500, David Abrahams wrote:\n\n> Good work.  I don't know if this matters to you, but I think on a 32-bit\n> platform you'll find that size_t, which is supposed to be able to hold\n> the size of the largest representable *memory block*, is only 4 bytes\n> large:\n\nThat should be fine; 32-bit systems can't deal with such large files\nanyway, since we want to address the whole thing. Getting around that\nwould, as Andreas mentioned, involve dealing with large files in chunks,\nsomething that would make the code a lot more complex.\n\nSo I think the answer is \"tough, if you want files >4G get a 64-bit\nmachine\". Which is unreasonable for a file system to say, but I think is\nfine for git.\n\n-Peff\n"}]}