Researchers present R2-OPD, a method addressing misalignment between teacher-derived rewards and actual reasoning progress in on-policy distillation for language models. The approach constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and one from independently estimated progress reward, selectively suppressing distillation signals when the rankings conflict. Experimental results show consistent improvements over standard on-policy distillation, particularly on reasoning-focused benchmarks.