North Micro Vision Instruct
North Micro Vision Instruct is a compact open-weight vision-language model for captioning, visual questions, grounding, and document understanding. It combines a 2B language model with a 400M vision encoder and preserves image aspect ratios through native-resolution processing. The language backbone supports 128K tokens, but multimodal training and validated operating conditions cover only 8K; this page uses the conservative multimodal limit. Apache-licensed weights make it a foundation for domain-specific fine-tuning and local experimentation. It supports interleaved images and text, multilingual prompts, and text output.
2026-08-12
2.4B
Decoder-only Transformer with native-resolution vision encoder
Apache-2.0
Specifications
- Parameters
- 2.4B
- Architecture
- Decoder-only Transformer with native-resolution vision encoder
- License
- Apache-2.0
- Context Window
- 8,192 tokens
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- North
- API Access
- Not Available
- Chat Interface
- Not Available
- Multilingual Support
- Yes
- Variants
- North-Micro-Vision-Instruct
Capabilities & Limitations
- Capabilities
- visiongroundingdocument understandingmultilingualmulti-image input
- Known Limitations
- Multimodal contexts beyond 8K are not validatedRequires task-specific evaluation before productionDoes not generate images
- Notable Use Cases
- visual question answeringdomain-specific vision assistantsdocument analysis