Molmo
Molmo is an open-weight Ai2 multimodal language-model family. The initial family combines vision encoders and pretrained language backbones with human-annotated PixMo image data, avoiding reliance on proprietary vision-language distillation. It answers visual questions, counts objects, and grounds responses with pointing. Backbones and license details vary across sizes, while the released vision-language training pipeline is open. It produces natural-language explanations from visual evidence and supports research on grounded interaction. Images or videos can be misinterpreted, particularly at fine resolution or across difficult temporal sequences.
2024-09-25
1B active MoE, 7B, 72B
Transformer language decoder with vision encoder
Apache-2.0 for documented checkpoint; backbone licenses may vary
Specifications
- Parameters
- 1B active MoE, 7B, 72B
- Architecture
- Transformer language decoder with vision encoder
- License
- Apache-2.0 for documented checkpoint; backbone licenses may vary
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- Molmo
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- MolmoE-1B7B-D7B-O72B
Capabilities & Limitations
- Capabilities
- visual question answeringspatial groundingpointing
- Known Limitations
- Perceptual responses may be inaccurateBackbone and serving support vary by checkpoint
- Notable Use Cases
- visual assistantsdocument/image understandingmultimodal research