A new study investigates why reinforcement learning with verifiable rewards (RLVR) improves single-sample accuracy while narrowing a language model’s solution diversity, reducing the benefit of test-time scaling. Using the Countdown task, whose solution space can be exhaustively enumerated, the researchers found that solution coverage drops by up to 67% during RLVR training, and that this contraction is concentrated at the first arithmetic operation chosen rather than in downstream reasoning steps. Guided by this finding, the authors show that late-layer parameter interpolation with earlier checkpoints increases solution coverage by 37% without any loss in accuracy.