From: Jeff King Date: Fri, 05 Dec 2025 18:01:06 GMT Subject: Re: [PATCH] packfile: skip decompressing and hashing blobs in add_promisor_object() Message-ID: <20251205180106.GC18566@coredump.intra.peff.net> In-Reply-To: <20251205174854.GA18566@coredump.intra.peff.net> On Fri, Dec 05, 2025 at 12:48:54PM -0500, Jeff King wrote: > OK, so we are checking the type up front and then skipping > parse_object() if we can. But there is already some logic inside > parse_object() for these kinds of optimizations. If we tell it we are > not interested in checking the hash of the objects, then it knows it can > skip loading the blob entirely. > > But it can _also_ use that flag for other things, like using the > commit-graph rather than loading individual commit objects. So doing > this: > > diff --git a/packfile.c b/packfile.c > index 9cc11b6dc5..01b992a4e1 100644 > --- a/packfile.c > +++ b/packfile.c > @@ -2310,7 +2310,8 @@ static int add_promisor_object(const struct object_id *oid, > we_parsed_object = 0; > } else { > we_parsed_object = 1; > - obj = parse_object(pack->repo, oid); > + obj = parse_object_with_flags(pack->repo, oid, > + PARSE_OBJECT_SKIP_HASH_CHECK); > } > > if (!obj) > > drops my linux.git case down to 49s. It's skipping the blobs (with no > need for your patch) and loading the commits out of the graph file. Note > that you may need to "git commit-graph write --reachable" to see the > effect (I think we do generate graphs by default in git-gc these days, > but I'm not sure if we do so right after cloning). Oh, and obviously it is skipping the hash computation on the objects, too. That's probably not as important as avoiding the object loads in the first place, but it may also be making a measurable difference on the ones we do load (notably trees here). -Peff