The paper proposes a unified four-coefficient parameterization for on-policy knowledge distillation that generalizes existing per-token gating methods by combining forward and reverse KL divergence losses across multiple channels with bias terms. It extends beyond prior single-channel approaches such as EOPD and ToDi, framing itself as a shared coordinate system for comparing gating designs. Tested on emotion and hate-speech classification tasks with Qwen models, configurations within this expanded family generally outperform single-channel baselines, though targeted replications show smaller effect sizes than originally reported.
