From: Patrick Steinhardt Date: Fri, 11 Sep 2026 12:23:46 GMT Subject: Re: [PATCH v2 10/10] builtin/fsck: move loose object verification into the loose source Message-ID: In-Reply-To: <875x0cnio5.fsf@emacs.iotcl.com> On Fri, Sep 11, 2026 at 01:15:06PM +0200, Toon Claes wrote: > Patrick Steinhardt writes: > > diff --git a/odb.h b/odb.h > > index 0bf6c8d7d2..b87f281cbd 100644 > > --- a/odb.h > > +++ b/odb.h > > @@ -218,6 +218,9 @@ enum odb_fsck_flags { > > > > /* Display a progress meter, if sensible. */ > > ODB_FSCK_PROGRESS = (1 << 1), > > + > > + /* Be extra verbose when checking the database. */ > > + ODB_FSCK_VERBOSE = (1 << 2), > > Shall we document this one is mutually exclusive with ODB_FSCK_PROGRESS? But is it really? Sure, we'll potentially have interleaving output where we print log messages followed by progress output. But as far as I can see, we have nothing where we fully interleave so that the progress output would be mangled. > > diff --git a/odb/source-loose.c b/odb/source-loose.c > > index f68d3c4d6c..efef9ca61f 100644 > > --- a/odb/source-loose.c > > +++ b/odb/source-loose.c > > @@ -1031,12 +1032,96 @@ static void odb_source_loose_free(struct odb_source *source) > > free(loose); > > } > > > > -static int odb_source_loose_fsck(struct odb_source *source UNUSED, > > - struct odb_fsck_options *opts UNUSED) > > +struct fsck_loose_data { > > + struct odb_source_loose *source; > > + struct odb_fsck_options *opts; > > + struct progress *progress; > > + bool error_found; > > +}; > > + > > +static int fsck_loose(const struct object_id *oid, const char *path, > > + void *cb_data) > > { > > + struct fsck_loose_data *data = cb_data; > > + enum object_type type = OBJ_NONE; > > + size_t size; > > + void *contents = NULL; > > + int eaten = 0; > > + struct object_info oi = OBJECT_INFO_INIT; > > + struct object_id real_oid = *null_oid(data->source->base.odb->repo->hash_algo); > > + int err = 0; > > + > > + oi.sizep = &size; > > + oi.typep = &type; > > + > > + if (read_loose_object(data->source->base.odb->repo, > > + path, oid, &real_oid, &contents, &oi) < 0) { > > + if (contents && !oideq(&real_oid, oid)) > > + err = error(_("%s: hash-path mismatch, found at: %s"), > > + oid_to_hex(&real_oid), path); > > + else > > + err = error(_("%s: object corrupt or missing: %s"), > > + oid_to_hex(oid), path); > > + } > > + if (err < 0) > > + goto out; > > + > > + if (!contents && type != OBJ_BLOB) > > + BUG("read_loose_object streamed a non-blob"); > > + > > + if (data->opts->object_cb(oid, type, size, contents, &eaten, > > + data->opts->object_payload)) { > > Should we guard data->opts->object_cb being NULL? I don't see a reason for that -- we don't currently have any callers that do, and we can still introduce this check if we ever grow one. Thanks! Patrick