Allen Institute for AI logo

Molmo

Allen Institute for AIOpen WeightsPending Human Review

Molmo is an open-weight Ai2 multimodal language-model family. The initial family combines vision encoders and pretrained language backbones with human-annotated PixMo image data, avoiding reliance on proprietary vision-language distillation. It answers visual questions, counts objects, and grounds responses with pointing. Backbones and license details vary across sizes, while the released vision-language training pipeline is open. It produces natural-language explanations from visual evidence and supports research on grounded interaction. Images or videos can be misinterpreted, particularly at fine resolution or across difficult temporal sequences.

2024-09-25
1B active MoE, 7B, 72B
Transformer language decoder with vision encoder
Apache-2.0 for documented checkpoint; backbone licenses may vary

Specifications

Parameters
1B active MoE, 7B, 72B
Architecture
Transformer language decoder with vision encoder
License
Apache-2.0 for documented checkpoint; backbone licenses may vary
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
Molmo
API Access
Not Available
Chat Interface
Not Available
Variants
MolmoE-1B7B-D7B-O72B

Capabilities & Limitations

Capabilities
visual question answeringspatial groundingpointing
Known Limitations
Perceptual responses may be inaccurateBackbone and serving support vary by checkpoint
Notable Use Cases
visual assistantsdocument/image understandingmultimodal research

Related Models