git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Memory issue with fast-import, why track branches?

From
JCJohn Chapman <thestar@fussycoder.id.au>
Date
Dec 21, 2008, 08:10 UTC
Message-ID
<1229847042.798.5.camel@therock.nsw.bigpond.net.au>
In-Reply-To
<94a0d4530812202154l26dfe0dfm49397c63dbfdfdf9@mail.gmail.com>

My first response was along the lines of "Why the heck are you storing sha1's like that!?", until I realised that you're not storing actual git sha1's, but mtn's hashes, which does make sense.

I'm doing something very similar with my perforce scripts, however I am doing a bit more magic instead of making so many branches.

Instead of making branches, I make a tag instead, for each and every changeset. Every time I make a new git commit, if I need to do it from a tag, I first read the tag and determine the sha1 I should use, and use that instead.

Alternatively, you could choose to manage your mapping yourself, and write them to a .git/mtg-git-map file.

On Sun, 2008-12-21 at 07:54 +0200, Felipe Contreras wrote:
Show 37 quoted lines
> Hi,
> 
> I tracked down an issue I have when importing a big repository. For
> some reason memory usage keeps increasing until there is no more
> memory.
> 
> Here is what valgrind shows:
> ==21034== 471,080,280 bytes in 114,517 blocks are still reachable in
> loss record 8 of 8
> ==21034==    at 0x4004BA2: calloc (vg_replace_malloc.c:397)
> ==21034==    by 0x806A340: xcalloc (wrapper.c:75)
> ==21034==    by 0x8063BC1: use_pack (sha1_file.c:808)
> ==21034==    by 0x8063DA9: unpack_object_header (sha1_file.c:1443)
> ==21034==    by 0x8064F4F: unpack_entry (sha1_file.c:1736)
> ==21034==    by 0x8065393: cache_or_unpack_entry (sha1_file.c:1606)
> ==21034==    by 0x8065464: read_packed_sha1 (sha1_file.c:2000)
> ==21034==    by 0x80655E5: read_object (sha1_file.c:2090)
> ==21034==    by 0x8065677: read_sha1_file (sha1_file.c:2106)
> ==21034==    by 0x8056AE9: parse_object (object.c:190)
> ==21034==    by 0x805E90A: write_ref_sha1 (refs.c:1214)
> ==21034==    by 0x804CC4F: update_branch (fast-import.c:1558)
> 
> After looking at the code my guess is that I have a humongous amount
> of branches.
> 
> Actually they are not really branches, but refs. For each git commit
> there's an original mtn ref that I store in 'refs/mtn/sha1', but since
> I'm using 'commit refs/mtn/sha1' to store it, a branch is created for
> every commit.
> 
> I guess there are many ways to fix the issue, but for starters I
> wonder why is fast-import keeping track of all the branches? In my
> case I would like fast-import to work exactly the same if I specify
> branches or not (I'll update them later).
> 
> Cheers.
> 
Previous: Felipe ContrerasNext: Felipe Contreras
Message 2 of 5 in “Memory issue with fast-import, why track branches?”
  1. Felipe ContrerasDec 21, 2008
  2. John ChapmanDec 21, 2008
  3. Felipe ContrerasDec 21, 2008
  4. Shawn O. PearceDec 21, 2008
  5. Felipe ContrerasDec 22, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.