Allen Institute for AI logo

MolmoPoint

Allen Institute for AIOpen WeightsPending Human Review

MolmoPoint is Ai2’s vision-language release with a token-based pointing mechanism that selects image regions directly from visual features rather than spelling out coordinates as ordinary text. It combines natural-language image and video understanding with spatial grounding, video pointing, and tracking. The 8B model uses Qwen3-8B and SigLIP; related 4B video and 8B GUI checkpoints specialize the same architecture. Weights, training code, and data support research on reliable visual interaction. Available implementations differ in which grounding functions they expose, and visual selections can still be inaccurate.

2026-03-18
4B, 8B language backbones
Transformer language decoder with SigLIP vision encoder and visual pointing mechanism
Apache-2.0

Specifications

Parameters
4B, 8B language backbones
Architecture
Transformer language decoder with SigLIP vision encoder and visual pointing mechanism
License
Apache-2.0
Type
multimodal
Modalities
textimagevideo

Benchmark Scores

Advanced Specifications

Model Family
Molmo
API Access
Not Available
Chat Interface
Not Available
Variants
8B4B Video8B GUI

Capabilities & Limitations

Capabilities
visual question answeringspatial groundingvideo trackingpointing
Known Limitations
Implementation support varies for pointing and trackingVisual grounding may be inaccurate
Notable Use Cases
grounded visual assistantsvideo analysisGUI research

Related Models