Researchers propose TurnSight, a framework for multi-turn tool-integrated reasoning that addresses the credit-assignment problem by deriving training supervision from an agent’s own execution outcomes rather than external references. The method aggregates token-level hindsight signals into turn-level assessments, selects reliable supervision through multi-horizon consensus, and modulates reinforcement-learning advantages while preserving their optimization direction. Across three benchmarks, TurnSight achieves state-of-the-art performance, including a 7.7% improvement over prior methods on an 8B model, while retaining full on-policy exploration capability.
