This paper identifies a ‘capability integration gap’ in multi-teacher on-policy distillation, where a controlled benchmark shows standard methods recover only 35.6% of the performance headroom available relative to an oracle domain-routed ensemble, traced to token-level optimization budget misallocation rather than gradient conflict. The authors introduce Open-MOPD, which adds token-share balancing, gap-aware dynamic budget allocation, and student reward refresh to correct the imbalance. These changes raise headroom recovery from 35.6% to 83.4%, and the team has fully open-sourced their training recipe, trajectories, and evaluation suite.
