Allen Institute for AI logo

Molmo 2

Allen Institute for AIOpen WeightsPending Human Review

Molmo 2 is an open-weight Ai2 multimodal language-model family. Molmo 2 extends the family to video reasoning, pointing, and tracking with open training data and code. The 8B checkpoint uses Qwen3-8B with SigLIP 2; a 7B Olmo-backed tier and smaller 4B variants share the release’s vision-language approach. Later ER and VideoPoint specializations remain related variants. It produces natural-language explanations from visual evidence and supports research on grounded interaction. Images or videos can be misinterpreted, particularly at fine resolution or across difficult temporal sequences.

2025-12-11
4B, 7B, 8B language backbones
Transformer language decoder with vision encoder
Apache-2.0 for documented checkpoint; backbone licenses may vary

Specifications

Parameters
4B, 7B, 8B language backbones
Architecture
Transformer language decoder with vision encoder
License
Apache-2.0 for documented checkpoint; backbone licenses may vary
Type
multimodal
Modalities
textimagevideo

Benchmark Scores

Advanced Specifications

Model Family
Molmo
API Access
Not Available
Chat Interface
Not Available
Variants
4B8BOlmo-backed 7BVideoPoint 4BER specialization

Capabilities & Limitations

Capabilities
visual question answeringspatial groundingpointingvideo understandingtracking
Known Limitations
Perceptual responses may be inaccurateBackbone and serving support vary by checkpoint
Notable Use Cases
visual assistantsdocument/image understandingmultimodal research

Related Models