Allen Institute for AI logo

MolmoWeb

Allen Institute for AIOpen WeightsPending Human Review

MolmoWeb is a released vision-language model trained for browser tasks from screenshots and natural-language instructions. Its 4B and 8B weights are distinct from the browser harness: the checkpoints generate textual reasoning, browser actions, and task-completion responses. Training uses MolmoWebMix, an open collection of synthetic and human browsing trajectories. Native and Transformers-compatible checkpoints are serving variants; compatibility fixes after launch do not establish a new model generation. Completing a web task depends on the external browser, observation/action interface, and execution permissions.

2026-03-24
4B, 8B
Vision-language Transformer
Apache-2.0

Specifications

Parameters
4B, 8B
Architecture
Vision-language Transformer
License
Apache-2.0
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
Molmo
API Access
Not Available
Chat Interface
Not Available
Variants
4B8BNative checkpointTransformers-compatible checkpoint

Capabilities & Limitations

Capabilities
browser reasoningvisual tool usetask completion
Known Limitations
Requires browser harnessDynamic websites and action errors can derail tasks
Notable Use Cases
browser-agent researchweb workflow automation
Tool Use Support
Yes

Related Models