This paper investigates on-policy distillation (OPD) by studying what happens when training uses minimal data, down to a single query, finding that one-shot OPD keeps improving for hundreds of training steps while recovering most of the gains of full-dataset training. The researchers describe OPD as data-overfed but algorithm-starved, showing through state-coverage analysis that a single query’s rollouts reach 71.5% of the states visited during full-data training, and that just 16 semantically diverse queries achieve near-full-data performance. The findings suggest future OPD work should prioritize algorithmic efficiency over expanding training datasets.