Re: [PATCH 2/2] packfile: recover when a multi-pack-index names a removed pack
- From
Jeff King <peff@peff.net>
- Date
- Aug 24, 2026, 04:48 UTC
- Message-ID
- <20260824044822.GA142844@coredump.intra.peff.net>
- In-Reply-To
- <CABPp-BEBbdmE9q+98gWq-wLzDdhJOyazcHF=pP95o5AcmgCv1Q@mail.gmail.com>
On Thu, Aug 20, 2026 at 06:36:09PM -0700, Elijah Newren wrote:
Show 30 quoted lines
> > > The false negative is not limited to one caller. Any reader > > > (cat-file, rev-list, pack-objects, ...) can spuriously fail with > > > "unable to read object", and callers that only ask whether an object > > > exists get a wrong answer too, since the OBJECT_INFO_QUICK path never > > > retries. Writers that merge in-core, such as "git replay", are hit > > > hardest: merge-ort treats the unreadable tree as a premature abort, sets > > > result.clean < 0, and returns without a result tree. > > > > Hm. Isn't there a slight variant of the race though for any caller that > > does not use OBJECT_INFO_QUICK? > > > > Namely, the packfile containing our object disappears and is being > > written to a new packfile, and that file is the only one containing it. > > Without OBJECT_INFO_QUICK we would be fine: we notice the object could > > not be found, and then we perform a second read that makes the "packed" > > backend reload its packfiles. It would find the new packfile, and > > because it's not covered by its MIDX it would use it to surface the > > object. But without OBJECT_INFO_QUICK that's not the case, as we would > > skip reloading packfiles altogether, and hence we would not be able to > > find that object at all. > > > > As far as I can see though, we don't seem to pass OBJECT_INFO_QUICK in > > any of the mentioned readers. I could very well be missing something > > here, but I would have thought that those readers are fine in this > > scenario? > > Nicely caught -- and you're right that the readers named above are > fine: they're all non-QUICK, so the second read reloads the packfiles > and finds the object in its new, non-MIDX-covered home, exactly as you > describe.
OK, so do I understand correctly that you _can't_ get the "unable to read object" result that the commit message claims? I.e., the reprepare / packfile reload is helps us (just like it does for the non-midx case when an idx has been mapped but the pack disappears before we open it).
So there is no bug there for non-QUICK callers. But then...
Show 6 quoted lines
> But the variant you describe is a real bug for QUICK callers that > don't get that second read -- e.g. upload-pack's object-existence > checks and mktree --batch. I have three more race-condition patches > to clean up and submit, and this is one of them: it forces the reload > even under OBJECT_INFO_QUICK once we notice a pack has vanished out > from under us.
This seems wrong. The whole point of the QUICK flag is that the caller is OK producing a false negative for an object lookup, and it would prefer that outcome to spending the time to reload. If there are callers passing QUICK that aren't OK with false negatives, they are broken and the fix should be there. But repreparing the packs for a QUICK miss is going to reintroduce the performance problems that QUICK was introduced to help.
So between the two cases, it sounds like things (or at least the low-level lookups) are working as designed, and there is no bug. Or am I misunderstanding something?
Show 5 quoted lines
> Your wording also makes me realize that my fix in this unsubmitted > patch still has a hole: it triggers when opening the pack .idx fails, > but if the timing is such that the .idx is already mmapped and only > the .pack has gone missing, it won't fire. I'll look into that before > submitting...and then clean up/submit my two other race fixes as well.
I think it would be fine, for the same reason that regular idx lookups are fine. In packfile_fill_entry() we call is_pack_valid(), checking that the pack is still there (and relying on its side effect of leaving the fd/mmap open so that it remains accessible even if the file is deleted).
-Peff