MolmoPoint
MolmoPoint is Ai2’s vision-language release with a token-based pointing mechanism that selects image regions directly from visual features rather than spelling out coordinates as ordinary text. It combines natural-language image and video understanding with spatial grounding, video pointing, and tracking. The 8B model uses Qwen3-8B and SigLIP; related 4B video and 8B GUI checkpoints specialize the same architecture. Weights, training code, and data support research on reliable visual interaction. Available implementations differ in which grounding functions they expose, and visual selections can still be inaccurate.
2026-03-18
4B, 8B language backbones
Transformer language decoder with SigLIP vision encoder and visual pointing mechanism
Apache-2.0
Specifications
- Parameters
- 4B, 8B language backbones
- Architecture
- Transformer language decoder with SigLIP vision encoder and visual pointing mechanism
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimagevideo
Benchmark Scores
Advanced Specifications
- Model Family
- Molmo
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- 8B4B Video8B GUI
Capabilities & Limitations
- Capabilities
- visual question answeringspatial groundingvideo trackingpointing
- Known Limitations
- Implementation support varies for pointing and trackingVisual grounding may be inaccurate
- Notable Use Cases
- grounded visual assistantsvideo analysisGUI research