git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] git-p4: improve performance with large files

From
Sam Hocevar <sam@zoy.org>
Date
Mar 5, 2009, 17:23 UTC
Message-ID
<20090305172332.GF25693@zoy.org>
In-Reply-To
<20090305100527.shmtfbdvk0ggsk4s@webmail.fussycoder.id.au>
On Thu, Mar 05, 2009, thestar@fussycoder.id.au wrote:
Show 11 quoted lines
> >   The current git-p4 way of concatenating strings performs in O(n^2)
> >and is therefore terribly slow with large files because of unnecessary
> >memory copies. The following patch makes the operation O(n).
> 
> The reason why it uses simple concatenation is to cut down on memory usage.
>  - It is a tradeoff.
> 
> I think the modification you have made below is reasonable, however be  
> aware that memory usage could double, which substantially reduce the  
> size of the changesets that git-p4 would be able to import /at all/,  
> rather than to merely be slow.
   Uhm, no. The memory usage could be an additional X, where X is the
size of the biggest file in the commit. Remember that commit() stores
the complete commit data in memory before sending it to fast-import.
Also, on my machine the extra memory is already used because at some
point, "text += foo" calls realloc() anyway and often duplicates the
memory used by text.
   The ideal solution is to use a generator and refactor the commit
handling as a stream. I am working on that but it involves deeper
changes, so as I am not sure it will be accepted, I'm providing the
attached compromise patch first. At least it solves the appaling speed
issue. I tuned it so that it never uses more than 32 MiB extra memory.
Signed-off-by: Sam Hocevar <sam@zoy.org>
---
 contrib/fast-import/git-p4 |   10 +++++++++-
 1 files changed, 9 insertions(+), 1 deletions(-)
diff --git a/contrib/fast-import/git-p4 b/contrib/fast-import/git-p4
index 3832f60..151ae1c 100755
--- a/contrib/fast-import/git-p4
+++ b/contrib/fast-import/git-p4
@@ -984,11 +984,19 @@ class P4Sync(Command):
         while j < len(filedata):
             stat = filedata[j]
             j += 1
+            data = []
             text = ''
             while j < len(filedata) and filedata[j]['code'] in ('text', 'unicod
e', 'binary'):
-                text += filedata[j]['data']
+                data.append(filedata[j]['data'])
                 del filedata[j]['data']
+                # p4 sends 4k chunks, make sure we don't use more than 32 MiB
+                # of additional memory while rebuilding the file data.
+                if len(data) > 8192:
+                    text += ''.join(data)
+                    data = []
                 j += 1
+            text += ''.join(data)
+            del data

             if not stat.has_key('depotFile'):
                 sys.stderr.write("p4 print fails with: %s\n" % repr(stat))
-- 
Sam.
Previous: thestar@fussycoder.id.auNext: thestar@fussycoder.id.au
Message 3 of 10 in “git-p4: improve performance with large files”
  1. git-p4: improve performance with large filesSam Hocevar, Mar 4, 2009
  2. thestar@fussycoder.id.auMar 4, 2009
  3. Sam HocevarMar 5, 2009
  4. thestar@fussycoder.id.auMar 6, 2009
  5. Junio C HamanoMar 6, 2009
  6. Han-Wen NienhuysMar 6, 2009
  7. Sam HocevarMar 6, 2009
  8. Junio C HamanoMar 6, 2009
  9. git-p4: improve performance with large filesSam Hocevar, Mar 6, 2009
  10. git-p4: improve performance when importing huge files by reducing the number of string concatenations while constraining memory usage.Sam Hocevar, Mar 7, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.