Molmo 2
Molmo 2 is an open-weight Ai2 multimodal language-model family. Molmo 2 extends the family to video reasoning, pointing, and tracking with open training data and code. The 8B checkpoint uses Qwen3-8B with SigLIP 2; a 7B Olmo-backed tier and smaller 4B variants share the release’s vision-language approach. Later ER and VideoPoint specializations remain related variants. It produces natural-language explanations from visual evidence and supports research on grounded interaction. Images or videos can be misinterpreted, particularly at fine resolution or across difficult temporal sequences.
2025-12-11
4B, 7B, 8B language backbones
Transformer language decoder with vision encoder
Apache-2.0 for documented checkpoint; backbone licenses may vary
Specifications
- Parameters
- 4B, 7B, 8B language backbones
- Architecture
- Transformer language decoder with vision encoder
- License
- Apache-2.0 for documented checkpoint; backbone licenses may vary
- Type
- multimodal
- Modalities
- textimagevideo
Benchmark Scores
Advanced Specifications
- Model Family
- Molmo
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- 4B8BOlmo-backed 7BVideoPoint 4BER specialization
Capabilities & Limitations
- Capabilities
- visual question answeringspatial groundingpointingvideo understandingtracking
- Known Limitations
- Perceptual responses may be inaccurateBackbone and serving support vary by checkpoint
- Notable Use Cases
- visual assistantsdocument/image understandingmultimodal research