WeiboAI researchers introduced CLR (Claim-Level Reliability Assessment), a training-free framework that improves LLM reasoning accuracy by directing verification effort at individual decision-critical claims rather than generating additional full solutions. The approach exploits the asymmetry that constructing a correct solution requires flawless reasoning throughout, while disproving an incorrect one requires finding just one flaw. On the CMIMC25 benchmark, CLR raised accuracy from 77.50% to 82.19% while using 37% fewer tokens than self-consistency approaches, and improved GPT-OSS-20B’s pass@1 by more than 27 percentage points.
