This paper finds that on-policy distillation’s teacher supervision contains substantial noise that doesn’t meaningfully affect student performance, and proposes On-Policy Self-Adaptation (OPSA), a supervision-free alternative using entropy-adaptive negative advantages that suppresses low-probability tokens and redistributes probability mass toward higher-probability ones. Tested on mathematical reasoning benchmarks, OPSA outperforms both baseline models and standard on-policy distillation by substantial margins.