Cohere logo

North Micro Vision Instruct

CohereOpen WeightsPending Human Review

North Micro Vision Instruct is a compact open-weight vision-language model for captioning, visual questions, grounding, and document understanding. It combines a 2B language model with a 400M vision encoder and preserves image aspect ratios through native-resolution processing. The language backbone supports 128K tokens, but multimodal training and validated operating conditions cover only 8K; this page uses the conservative multimodal limit. Apache-licensed weights make it a foundation for domain-specific fine-tuning and local experimentation. It supports interleaved images and text, multilingual prompts, and text output.

2026-08-12
2.4B
Decoder-only Transformer with native-resolution vision encoder
Apache-2.0

Specifications

Parameters
2.4B
Architecture
Decoder-only Transformer with native-resolution vision encoder
License
Apache-2.0
Context Window
8,192 tokens
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
North
API Access
Not Available
Chat Interface
Not Available
Multilingual Support
Yes
Variants
North-Micro-Vision-Instruct

Capabilities & Limitations

Capabilities
visiongroundingdocument understandingmultilingualmulti-image input
Known Limitations
Multimodal contexts beyond 8K are not validatedRequires task-specific evaluation before productionDoes not generate images
Notable Use Cases
visual question answeringdomain-specific vision assistantsdocument analysis

Related Models