Tencent researchers describe WeMM-Embedding, a family of multimodal embedding models in 2B, 4B, and 9B parameter sizes trained through large-scale alignment followed by refinement with curated data and cross-scale knowledge transfer. The models support text, images, videos, visual documents, and arbitrarily interleaved multimodal inputs with adjustable output dimensions. The 2B variant outperforms the previous leading 8B open-source baseline on MMEB-v2, while the 9B variant reaches a state-of-the-art score of 80.6 and has been deployed across multiple WeChat platform applications.
