NeoMME introduces a family of compact 260M and 800M-parameter bidirectional encoders that process multilingual text and raw image patches within a single transformer, avoiding the overhead of repurposing full vision-language generative models for retrieval tasks. Pretrained from scratch with a masked discrete-diffusion objective and fine-tuned for visual document retrieval, the smaller NeoMME-260M outperforms all sub-800M models on the ViDoRe v3 benchmark while encoding pages at roughly twice the throughput of ColModernVBERT, and the team has released the models publicly on Hugging Face.
