This paper formalizes what is required to legitimately ‘replay’ a claim attached to an LLM evaluation metric, then applies that framework to audit all 124 mechanically eligible units in the Inspect Evals suite at a pinned commit. It finds that 110 of the 124 evaluation units stop before producing a deterministic result because required historical evidence or semantic grounding is unavailable, undermining confidence in point-in-time benchmark comparisons. The audit produces typed dispositions for each unit rather than a single blanket robust/not-robust verdict.
