git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v2] packfile: skip decompressing and hashing blobs in add_promisor_object()

From
Aaron Plattner <aplattner@nvidia.com>
Date
Dec 6, 2025, 00:20 UTC
Message-ID
<20251206002014.2066644-1-aplattner@nvidia.com>

When is_promisor_object() is called for the first time, it lazily initializes a set of all promisor objects by iterating through all objects in promisor packs. For each object, add_promisor_object() calls parse_object(), which decompresses and hashes the entire object.

For repositories with large pack files, this can take an extremely long time. For example, on a production repository with a 176 GB promisor pack:

 $ time ~/git/git/git-rev-list --objects --all --exclude-promisor-objects --quiet
 ________________________________________________________
 Executed in   76.10 mins    fish           external
    usr time   72.10 mins    1.83 millis   72.10 mins
    sys time    3.56 mins    0.17 millis    3.56 mins

add_promisor_object() needs the full object for trees, commits, and tags. But blobs contain no references to other objects, so the function can just insert their oids into the set and move on.

parse_object_with_flags() has code to skip decompressing blobs, but it unfortunately doesn't work with the objects created by mark_uninteresting() because they have obj->type == OBJ_NONE. Update parse_object_with_flags() to handle blobs and trees that are in this state, and then update add_promisor_object() to use PARSE_OBJECT_SKIP_HASH_CHECK.

This improves performance for very large pack files significantly:
 $ time ~/git/git/git-rev-list --objects --all --exclude-promisor-objects --quiet
 ________________________________________________________
 Executed in  117.63 secs    fish           external
    usr time   45.56 secs    1.09 millis   45.56 secs
    sys time   37.91 secs    1.05 millis   37.91 secs
Signed-off-by: Aaron Plattner <aplattner@nvidia.com>
---
v2: Fix PARSE_OBJECT_SKIP_HASH_CHECK with UNINTERESTING objects, use it
in parse_object_with_flags.
 object.c   | 4 ++--
 packfile.c | 3 ++-
 2 files changed, 4 insertions(+), 3 deletions(-)
diff --git a/object.c b/object.c
index b08fc7a163..4669b8d65e 100644
--- a/object.c
+++ b/object.c
@@ -328,7 +328,7 @@ struct object *parse_object_with_flags(struct repository *r,
 			return &commit->object;
 	}
 
-	if ((!obj || obj->type == OBJ_BLOB) &&
+	if ((!obj || obj->type == OBJ_NONE || obj->type == OBJ_BLOB) &&
 	    odb_read_object_info(r->objects, oid, NULL) == OBJ_BLOB) {
 		if (!skip_hash && stream_object_signature(r, repl) < 0) {
 			error(_("hash mismatch %s"), oid_to_hex(oid));
@@ -344,7 +344,7 @@ struct object *parse_object_with_flags(struct repository *r,
 	 * have the on-disk object with the correct type.
 	 */
 	if (skip_hash && discard_tree &&
-	    (!obj || obj->type == OBJ_TREE) &&
+	    (!obj || obj->type == OBJ_NONE || obj->type == OBJ_TREE) &&
 	    odb_read_object_info(r->objects, oid, NULL) == OBJ_TREE) {
 		return &lookup_tree(r, oid)->object;
 	}
diff --git a/packfile.c b/packfile.c
index 9cc11b6dc5..01b992a4e1 100644
--- a/packfile.c
+++ b/packfile.c
@@ -2310,7 +2310,8 @@ static int add_promisor_object(const struct object_id *oid,
 		we_parsed_object = 0;
 	} else {
 		we_parsed_object = 1;
-		obj = parse_object(pack->repo, oid);
+		obj = parse_object_with_flags(pack->repo, oid,
+					      PARSE_OBJECT_SKIP_HASH_CHECK);
 	}
 
 	if (!obj)
-- 
2.52.0
Next: Jeff King
Message 1 of 4 in “packfile: skip decompressing and hashing blobs in add_promisor_object()”
  1. packfile: skip decompressing and hashing blobs in add_promisor_object()Aaron Plattner, Dec 6, 2025
  2. Jeff KingDec 6, 2025
  3. Aaron PlattnerDec 6, 2025
  4. Jeff KingDec 8, 2025

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.