git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v4] convert: avoid high memory footprint

From
HGHaritha via GitGitGadget <gitgitgadget@gmail.com>
Date
Jul 26, 2024, 14:00 UTC
Message-ID
<pull.1744.v4.git.git.1722002432630.gitgitgadget@gmail.com>
In-Reply-To
<pull.1744.v3.git.git.1721975234873.gitgitgadget@gmail.com>
From: D Harithamma <harithamma.d@ibm.com>

When Git adds a file requiring encoding conversion and tracing of encoding conversion is not requested via the GIT_TRACE_WORKING_TREE_ENCODING environment variable, the `trace_encoding()` function still allocates & prepares "human readable" copies of the file contents before and after conversion to show in the trace. This results in a high memory footprint and increased runtime without providing any user-visible benefit.

This fix introduces an early exit from the `trace_encoding()` function when tracing is not requested, preventing unnecessary memory allocation and processing.

Signed-off-by: Harithamma D <harithamma.d@ibm.com>
---
    Fix to avoid high memory footprint
    
    This fix avoids high memory footprint when adding files that require
    conversion
    
    Git has a trace_encoding routine that prints trace output when
    GIT_TRACE_WORKING_TREE_ENCODING=1 is set. This environment variable is
    used to debug the encoding contents. When a 40MB file is added, it
    requests close to 1.8GB of storage from xrealloc which can lead to out
    of memory errors. However, the check for GIT_TRACE_WORKING_TREE_ENCODING
    is done after the string is allocated. This resolves high memory
    footprints even when GIT_TRACE_WORKING_TREE_ENCODING is not active. This
    fix adds an early exit to avoid the unnecessary memory allocation.
Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-1744%2FHarithaIBM%2FmemFootprintFix-v4
Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-1744/HarithaIBM/memFootprintFix-v4
Pull-Request: https://github.com/git/git/pull/1744
Range-diff vs v3:
 1:  d864de64380 ! 1:  50758a4fb94 Fix to avoid high memory footprint
     @@ Metadata
      Author: D Harithamma <harithamma.d@ibm.com>
      
       ## Commit message ##
     -    Fix to avoid high memory footprint
     +    convert: avoid high memory footprint
      
          When Git adds a file requiring encoding conversion and tracing of encoding
          conversion is not requested via the GIT_TRACE_WORKING_TREE_ENCODING
 convert.c | 3 +++
 1 file changed, 3 insertions(+)
diff --git a/convert.c b/convert.c
index d8737fe0f2d..c4ddc4de81b 100644
--- a/convert.c
+++ b/convert.c
@@ -324,6 +324,9 @@ static void trace_encoding(const char *context, const char *path,
 	struct strbuf trace = STRBUF_INIT;
 	int i;
 
+	if (!trace_want(&coe))
+		return;
+
 	strbuf_addf(&trace, "%s (%s, considered %s):\n", context, path, encoding);
 	for (i = 0; i < len && buf; ++i) {
 		strbuf_addf(

base-commit: 557ae147e6cdc9db121269b058c757ac5092f9c9
-- 
gitgitgadget
Previous: Torsten BögershausenNext: Haritha via GitGitGadget
Message 8 of 15 in “Fix to avoid high memory footprint”
  1. Fix to avoid high memory footprintHaritha via GitGitGadget, Jul 16, 2024
  2. Jeff KingJul 17, 2024
  3. Fix to avoid high memory footprintHaritha via GitGitGadget, Jul 24, 2024
  4. Junio C HamanoJul 24, 2024
  5. Jeff KingJul 24, 2024
  6. Fix to avoid high memory footprintHaritha via GitGitGadget, Jul 26, 2024
  7. Torsten BögershausenJul 26, 2024
  8. convert: avoid high memory footprintHaritha via GitGitGadget, Jul 26, 2024
  9. convert: return early when not tracingHaritha via GitGitGadget, Jul 30, 2024
  10. Junio C HamanoJul 31, 2024
  11. Haritha DJul 31, 2024
  12. convert: return early when not tracingHaritha via GitGitGadget, Jul 31, 2024
  13. Junio C HamanoJul 26, 2024
  14. Junio C HamanoJul 26, 2024
  15. Haritha DJul 30, 2024

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.