SMELT loops the middle half of layers twice in Mixture-of-Experts Transformers while holding FLOPs, parameter count, and KV cache constant. The authors derive scaling laws showing the approach achieves 6.8-18.0% training FLOP savings on the compute-optimal frontier. Mechanistic analysis across model sizes up to 54B parameters shows repeated layer visits reduce attention artifacts and redirect focus toward content-relevant tokens.
