Published July 30, 2026, this article presents mDenseOn and mLateOn, two multilingual retrieval models developed by LightOn AI researchers. The 307-million-parameter models were trained on a 2.8-billion-pair dataset built through a translate-train methodology that extends an English training recipe to eight additional languages. mLateOn, which uses a late-interaction architecture, outperformed its dense counterpart, reaching 57.56 NDCG@10 on the BEIR benchmark and showing notably better transfer to languages not included in training.
