Researchers from Kakao Corp introduced a compute-efficient hyperparameter transfer framework for large-scale Mixture-of-Experts models, adapting Maximal Update Parameterization (μP) to architectures using Multi-head Latent Attention and the Muon optimizer. The method shows optimal learning rates transfer consistently across model widths and extends that transferability along the token dimension via a predictive scaling law. This lets teams tune hyperparameters for trillion-token training runs without running prohibitively expensive full-scale sweeps.
