Liquid AI introduced LFM2.5-VL-3B, a vision-language model that pairs a SigLIP2 400M NaFlex vision encoder with the LFM2.5-2.6B text backbone. The model was pre-trained on roughly 34 trillion tokens with quadrupled vision data, then refined through supervised fine-tuning, knowledge distillation, and multi-reward reinforcement learning. Tuned for edge deployment, it reaches 228 tokens per second on a MacBook M5 Max while remaining competitive on benchmarks for document understanding, object grounding, screen UI recognition, and function calling.
