Cohere Labs has released North Micro Vision, a 2.4-billion-parameter vision-language model that pairs a custom 400-million-parameter native-resolution vision encoder with an in-house 2-billion-parameter language model. The model was trained across four stages that progressively increase input resolution up to 1654×2339 pixels while incorporating dense captions and OCR data. Its architecture injects patch embeddings from multiple encoder layers into early language-model layers to preserve fine-grained spatial detail needed for document and form processing. Cohere Labs reports particularly strong results on document and chart understanding benchmarks alongside competitive performance on multilingual and robustness evaluations.
