{"thread":{"id":"18204","subject":"Git memory usage (1): fast-import","startedAt":"2009-03-07T20:19:20Z","lastAt":"2009-03-07T21:01:30Z","messageCount":2,"participants":["Sam Hocevar","Shawn O. Pearce"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"107339","messageId":"20090307201920.GE12880@zoy.org","threadId":"18204","inReplyTo":null,"subject":"Git memory usage (1): fast-import","fromName":"Sam Hocevar","fromEmail":"sam@zoy.org","sentAt":"2009-03-07T20:19:20Z","receivedAt":"2009-03-07T20:19:20Z","isPatch":false,"sender":{"key":"sam@zoy.org","avatar":"https://gravatar.com/avatar/1fc1e5d8c3a8d737f14572135671adfbc0e61ffba5c3bc1d1b8a6f4aac764470?d=mp&s=160"},"body":"   I joined a project that uses very large binary files (up to 1 GiB) in\na p4 repository and as I would like to use Git, I am trying to make it\nmore memory-efficient when handling huge files.\n\n   The first problem I am hitting is with fast-import. Currently it\nkeeps the last imported file in memory (end of store_object()) in order\nto find interesting deltas with the next file. Since most huge binary\nfiles are already compressed, it means that fast-importing two large\nX MiB files is going to use at least 3X MiB of memory: once for the\nfirst file, once for the second file, and once for the deflate data\nthat is going to be as large as the file itself.\n\n   In practice, it takes even more memory than that. Experiment shows\nthat importing six 100 MiB files made of urandom data takes 370 MiB of\nmemory (http://zoy.org/~sam/git/git-memory-usage.png) (simple script\navailable at http://zoy.org/~sam/git/gencommit.txt). I am unable to plot\nhow it behaves with 1 GiB files since I don't have enough memory, but I\ndon't see why the trend wouldn't stand.\n\n   I can understand it not being a priority, but I'm trying to think of\nacceptable ways to fix this that do not mess with Git's performance in\nmore usual cases. Here is what I can think of:\n\n   - stop trying to compute deltas in fast-import and leave that task\n   to other tools (optionally, define a file size threshold beyond\n   which the last file is not kept in memory, and maybe make that a\n   configuration option).\n\n   - use a temporary file to store the deflate data when it reaches a\n   given size threshold (and maybe make that a configuration option).\n\n   - also, I haven't tracked all strbuf_* uses in fast-import, but I got\n   the feeling that strbuf_release() could be used in a few places\n   instead of strbuf_setlen(0) in order to free some memory.\n\n   Any thoughts?\n\n-- \nSam.\n"},{"id":"107345","messageId":"20090307210130.GO16213@spearce.org","threadId":"18204","inReplyTo":"20090307201920.GE12880@zoy.org","subject":"Re: Git memory usage (1): fast-import","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2009-03-07T21:01:30Z","receivedAt":"2009-03-07T21:01:30Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Sam Hocevar <sam@zoy.org> wrote:\n>    I joined a project that uses very large binary files (up to 1 GiB) in\n> a p4 repository and as I would like to use Git, I am trying to make it\n> more memory-efficient when handling huge files.\n\nYikes.  As you saw, this won't play well...\n \n>    In practice, it takes even more memory than that. Experiment shows\n> that importing six 100 MiB files made of urandom data takes 370 MiB of\n> memory [...]\n\nYes.\n\nAs you saw, this is the last object, the current object, the delta\nindex of the last object (in order to more efficiently compare the\ncurrent one to it), and the deflate buffer for the current object,\noh, and probably memory fragmentation....\n\nI'm not surprised a 100 MiB file turned into 370 MiB heap usage.\n \n>    - stop trying to compute deltas in fast-import and leave that task\n>    to other tools\n\nThis isn't practical for source code imports, unless we do...\n\n> (optionally, define a file size threshold beyond\n>    which the last file is not kept in memory, and maybe make that a\n>    configuration option).\n\nwhat you suggest here.  fast-import is faster than other methods\nbecause we get some delta compression on the content, so the output\npack uses up less virtual memory when the front-end or end-user\nfinally gets around to doing `git repack -a -d -f` to recompute\nthe delta chains.\n\n>    - use a temporary file to store the deflate data when it reaches a\n>    given size threshold (and maybe make that a configuration option).\n\nZoiks.  There's no reason for that.\n\nA better method would be to just look at the size of the incoming\nblob, and if its over some configured threshold (default e.g. 100\nMB is perhaps sane) we just stream the data through deflate()\nand into the pack file, with no delta compression.\n\nThat would also bypass the \"massive\" buffer in the last object slot,\nas you point out above.\n \n>    - also, I haven't tracked all strbuf_* uses in fast-import, but I got\n>    the feeling that strbuf_release() could be used in a few places\n>    instead of strbuf_setlen(0) in order to free some memory.\n\nExamples?  I haven't gone through the code in detail since it\nwas modified to use strbufs.  But I had the feeling that the code\nwasn't freeing strbufs that it would just reuse on the next command,\nand that are likely to be \"smallish\", e.g. just a few KiBs in size.\n\n-- \nShawn.\n"}]}