MolmoWeb
MolmoWeb is a released vision-language model trained for browser tasks from screenshots and natural-language instructions. Its 4B and 8B weights are distinct from the browser harness: the checkpoints generate textual reasoning, browser actions, and task-completion responses. Training uses MolmoWebMix, an open collection of synthetic and human browsing trajectories. Native and Transformers-compatible checkpoints are serving variants; compatibility fixes after launch do not establish a new model generation. Completing a web task depends on the external browser, observation/action interface, and execution permissions.
2026-03-24
4B, 8B
Vision-language Transformer
Apache-2.0
Specifications
- Parameters
- 4B, 8B
- Architecture
- Vision-language Transformer
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- Molmo
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- 4B8BNative checkpointTransformers-compatible checkpoint
Capabilities & Limitations
- Capabilities
- browser reasoningvisual tool usetask completion
- Known Limitations
- Requires browser harnessDynamic websites and action errors can derail tasks
- Notable Use Cases
- browser-agent researchweb workflow automation
- Tool Use Support
- Yes