{"thread":{"id":"40362","subject":"git-svn: cat-file memory usage","startedAt":"2015-09-16T11:00:48Z","lastAt":"2015-09-16T16:31:47Z","messageCount":4,"participants":["Victor Leschuk","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"270117","messageId":"6AE1604EE3EC5F4296C096518C6B77EE5D0FDAB9CB@mail.accesssoftek.com","threadId":"40362","inReplyTo":null,"subject":"git-svn: cat-file memory usage","fromName":"Victor Leschuk","fromEmail":"vleschuk@accesssoftek.com","sentAt":"2015-09-16T11:00:48Z","receivedAt":"2015-09-16T11:00:48Z","isPatch":false,"sender":{"key":"vleschuk@accesssoftek.com","avatar":null},"body":"Hello all,\n\nWe are currently getting acquainted with git-svn tool and have experienced few problems with it. The main issue is memory usage during \"git svn clone\": on large repositories the perl and git processes are using significant amount of memory. \n\nI have conducted several tests with different repositories. I have created mirrors of Trac project (http://trac.edgewall.org/ - rather small repo, ~14000 commits) and FreeBSD base repo (~280000 commits). Here is the summary of my tests (to eliminate network issues all clones were performed for file:// repos):\n\n * git svn clone of trac  takes about 1 hour \n * git svn clone of FreeBSD has already taken more than 3 days and still running (currently has cloned about 40% of revisions)\n * git cat-file process memory footprint keeps growing during the clone process (see figure attached)\n\nThe main issue here is git cat-file consuming memory. The attached figure is for small repository which takes about an hour to clone, however on my another machine where FreeBSD clone is currently running the git cat-file has already taken more than 1Gb of memory (RSS) and has overgrown the parent perl process (~300-400 Mb). \n\nI have valgrind'ed the git-cat-file (which is running is --batch mode during the whole clone) and found no serious leaks (about 100 bytes definitely leaked), so all memory is carefully freed, but the heap usage grows maybe due to fragmentation or smth else. When I looked through the code I found out that most of heap allocations are called from batch_object_write() function (strbuf_expand -> realloc).\n\nSo I have found two possible workarounds for the issue: \n\n * Set GIT_ALLOC_LIMIT variable - it does reduce the memory footprint but slows down the process\n * In perl code do not run git cat-file in batch mode (in Git::SVN::apply_textdelta) but rather run it as separate commands each time\n\n   my $size = $self->command_oneline('cat-file', '-s', $sha1);\n   # .....\n   my ($in, $c) = $self->command_output_pipe('cat-file', 'blob', $sha1);\n\nThe second approach doesn't slow down the whole process at all (~72 minutes to clone repo both with --batch mode and without).\n\nSo the question is: what would be the correct approach to fight the problem with cat-file memory usage: maybe we should get rid of batch mode in perl code, or somehow tune allocation policy in C code?\n\nPlease let me know your thoughts. \n\n--\nBest Regards,\nVictor Leschuk"},{"id":"270119","messageId":"20150916115642.GA5104@sigill.intra.peff.net","threadId":"40362","inReplyTo":"6AE1604EE3EC5F4296C096518C6B77EE5D0FDAB9CB@mail.accesssoftek.com","subject":"Re: git-svn: cat-file memory usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2015-09-16T11:56:42Z","receivedAt":"2015-09-16T11:56:42Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 16, 2015 at 04:00:48AM -0700, Victor Leschuk wrote:\n\n>  * git svn clone of trac  takes about 1 hour \n>  * git svn clone of FreeBSD has already taken more than 3 days and\n>  still running (currently has cloned about 40% of revisions)\n\nI haven't worked with git-svn in a long time, but I doubt that it is the\nfastest way to do a large repository import. You might want to look into\na tool like svn2git or reposurgeon to do the initial import.\n\n> I have valgrind'ed the git-cat-file (which is running is --batch mode\n> during the whole clone) and found no serious leaks (about 100 bytes\n> definitely leaked), so all memory is carefully freed, but the heap\n> usage grows maybe due to fragmentation or smth else. When I looked\n> through the code I found out that most of heap allocations are called\n> from batch_object_write() function (strbuf_expand -> realloc).\n\nCertainly we will call strbuf_expand once per object. I would have\nexpected we would call read_sha1_file(), too. It looks like we always\ntry to stream blobs, but I think we have to fall back to reading the\nwhole object if there are deltas.\n\nYou can try this patch, which will reuse the same strbuf over and over:\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 07baad1..73f338c 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -256,7 +256,7 @@ static void print_object_or_die(struct batch_options *opt, struct expand_data *d\n static void batch_object_write(const char *obj_name, struct batch_options *opt,\n \t\t\t       struct expand_data *data)\n {\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstatic struct strbuf buf = STRBUF_INIT;\n \n \tif (sha1_object_info_extended(data->sha1, &data->info, LOOKUP_REPLACE_OBJECT) < 0) {\n \t\tprintf(\"%s missing\\n\", obj_name ? obj_name : sha1_to_hex(data->sha1));\n@@ -264,10 +264,10 @@ static void batch_object_write(const char *obj_name, struct batch_options *opt,\n \t\treturn;\n \t}\n \n+\tstrbuf_reset(&buf);\n \tstrbuf_expand(&buf, opt->format, expand_format, data);\n \tstrbuf_addch(&buf, '\\n');\n \tbatch_write(opt, buf.buf, buf.len);\n-\tstrbuf_release(&buf);\n \n \tif (opt->print_contents) {\n \t\tprint_object_or_die(opt, data);\n\nThat will reduce your reallocs due to strbuf_expand, though I'm doubtful\nthat it will solve the problem (and if it does, I think the right\nsolution is probably to look into using a better allocator than what\nyour system malloc() is providing).\n\n>  * In perl code do not run git cat-file in batch mode (in\n>  Git::SVN::apply_textdelta) but rather run it as separate commands\n>  each time\n> \n>    my $size = $self->command_oneline('cat-file', '-s', $sha1);\n>    # .....\n>    my ($in, $c) = $self->command_output_pipe('cat-file', 'blob', $sha1);\n> \n> The second approach doesn't slow down the whole process at all (~72\n> minutes to clone repo both with --batch mode and without).\n\nI'm surprised the startup cost of the process doesn't make an impact,\nbut maybe it gets lost in the noise of the rest of the work (AFAICT, the\npoint of this cat-file is to retrieve a blob, apply a delta to it, and\nthen write out the resulting object; that write is probably a lot more\nexpensive).\n\n-Peff\n"},{"id":"270123","messageId":"6AE1604EE3EC5F4296C096518C6B77EE5D0FDAB9CD@mail.accesssoftek.com","threadId":"40362","inReplyTo":"20150916115642.GA5104@sigill.intra.peff.net","subject":"RE: git-svn: cat-file memory usage","fromName":"Victor Leschuk","fromEmail":"vleschuk@accesssoftek.com","sentAt":"2015-09-16T13:40:23Z","receivedAt":"2015-09-16T13:40:23Z","isPatch":false,"sender":{"key":"vleschuk@accesssoftek.com","avatar":null},"body":"Hello Jeff, thanks for the advice.\n\nUnfortunately using patch didn't change the situation. I will run some tests with alternate allocators (looking at jemalloc and tcmalloc). As for alternate tools: as far as I understood svn2git calls 'git svn' itself. So I assume it can't fix the memory usage or speed up clone process... Correct me if I'm wrong.\n\nReposurgeon looks interesting... Will give it a try. \n\nBtw, what do you think of getting rid of batch mode for clone/fetch in perl code. It really hardly has any impact on performance but reduces memory usage a lot.\n\n--\nBest Regards,\nVictor\n________________________________________\nFrom: Jeff King [peff@peff.net]\nSent: Wednesday, September 16, 2015 4:56 AM\nTo: Victor Leschuk\nCc: git@vger.kernel.org\nSubject: Re: git-svn: cat-file memory usage\n\nOn Wed, Sep 16, 2015 at 04:00:48AM -0700, Victor Leschuk wrote:\n\n>  * git svn clone of trac  takes about 1 hour\n>  * git svn clone of FreeBSD has already taken more than 3 days and\n>  still running (currently has cloned about 40% of revisions)\n\nI haven't worked with git-svn in a long time, but I doubt that it is the\nfastest way to do a large repository import. You might want to look into\na tool like svn2git or reposurgeon to do the initial import.\n\n> I have valgrind'ed the git-cat-file (which is running is --batch mode\n> during the whole clone) and found no serious leaks (about 100 bytes\n> definitely leaked), so all memory is carefully freed, but the heap\n> usage grows maybe due to fragmentation or smth else. When I looked\n> through the code I found out that most of heap allocations are called\n> from batch_object_write() function (strbuf_expand -> realloc).\n\nCertainly we will call strbuf_expand once per object. I would have\nexpected we would call read_sha1_file(), too. It looks like we always\ntry to stream blobs, but I think we have to fall back to reading the\nwhole object if there are deltas.\n\nYou can try this patch, which will reuse the same strbuf over and over:\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 07baad1..73f338c 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -256,7 +256,7 @@ static void print_object_or_die(struct batch_options *opt, struct expand_data *d\n static void batch_object_write(const char *obj_name, struct batch_options *opt,\n                               struct expand_data *data)\n {\n-       struct strbuf buf = STRBUF_INIT;\n+       static struct strbuf buf = STRBUF_INIT;\n\n        if (sha1_object_info_extended(data->sha1, &data->info, LOOKUP_REPLACE_OBJECT) < 0) {\n                printf(\"%s missing\\n\", obj_name ? obj_name : sha1_to_hex(data->sha1));\n@@ -264,10 +264,10 @@ static void batch_object_write(const char *obj_name, struct batch_options *opt,\n                return;\n        }\n\n+       strbuf_reset(&buf);\n        strbuf_expand(&buf, opt->format, expand_format, data);\n        strbuf_addch(&buf, '\\n');\n        batch_write(opt, buf.buf, buf.len);\n-       strbuf_release(&buf);\n\n        if (opt->print_contents) {\n                print_object_or_die(opt, data);\n\nThat will reduce your reallocs due to strbuf_expand, though I'm doubtful\nthat it will solve the problem (and if it does, I think the right\nsolution is probably to look into using a better allocator than what\nyour system malloc() is providing).\n\n>  * In perl code do not run git cat-file in batch mode (in\n>  Git::SVN::apply_textdelta) but rather run it as separate commands\n>  each time\n>\n>    my $size = $self->command_oneline('cat-file', '-s', $sha1);\n>    # .....\n>    my ($in, $c) = $self->command_output_pipe('cat-file', 'blob', $sha1);\n>\n> The second approach doesn't slow down the whole process at all (~72\n> minutes to clone repo both with --batch mode and without).\n\nI'm surprised the startup cost of the process doesn't make an impact,\nbut maybe it gets lost in the noise of the rest of the work (AFAICT, the\npoint of this cat-file is to retrieve a blob, apply a delta to it, and\nthen write out the resulting object; that write is probably a lot more\nexpensive).\n\n-Peff\n"},{"id":"270131","messageId":"20150916163146.GA28401@sigill.intra.peff.net","threadId":"40362","inReplyTo":"6AE1604EE3EC5F4296C096518C6B77EE5D0FDAB9CD@mail.accesssoftek.com","subject":"Re: git-svn: cat-file memory usage","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2015-09-16T16:31:47Z","receivedAt":"2015-09-16T16:31:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Sep 16, 2015 at 06:40:23AM -0700, Victor Leschuk wrote:\n\n> Unfortunately using patch didn't change the situation. I will run some\n> tests with alternate allocators (looking at jemalloc and tcmalloc). As\n> for alternate tools: as far as I understood svn2git calls 'git svn'\n> itself. So I assume it can't fix the memory usage or speed up clone\n> process... Correct me if I'm wrong.\n\nI think there are actually several tools calling themselves svn2git.\nThere was a C tool once upon a time, but it looks fairly inactive, and\nthe top search hit for svn2git does turn up a git-svn wrapper. Like I\nsaid, I am not very up on the current state of affairs.\n\n> Btw, what do you think of getting rid of batch mode for clone/fetch in\n> perl code. It really hardly has any impact on performance but reduces\n> memory usage a lot.\n\nI'd worry there are other cases where it does impact performance (e.g.,\nperhaps smaller blobs), but I don't know enough about the git-svn\ninternals to say much more.\n\n-Peff\n"}]}