Researchers propose Test-Time Policy Optimization (TTPO), a label-free training objective for adapting language models to math reasoning tasks at test time. TTPO distills rollouts that agree with a majority-vote pseudo-label while using reinforcement learning to penalize disagreeing rollouts, addressing the fragility of naive pseudo-label supervision. Without any ground-truth labels, TTPO matches label-supervised baselines on five competition-level benchmarks and raises Qwen3-1.7B accuracy from 38.0% to 45.2%.
