UI-Venus-1.5
UI-Venus-1.5 is a distinct vision-language model specialization for GUI interaction, combining screenshot understanding with natural-language task reasoning and structured actions. The 1.5 release offers 2B, 8B, and 30B-A3B checkpoints for grounding and end-to-end mobile agents, with Apache-2.0 weights. Local inference requires an external action parser and execution framework; the model itself emits language reasoning and actions.
2026-02-09
2B, 8B, 30B-A3B
Transformer vision-language model with dense and MoE sizes
Apache-2.0
Specifications
- Parameters
- 2B, 8B, 30B-A3B
- Architecture
- Transformer vision-language model with dense and MoE sizes
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- UI-Venus
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- UI-Venus-1.5-2BUI-Venus-1.5-8BUI-Venus-1.5-30B-A3B
Capabilities & Limitations
- Capabilities
- GUI groundingtask reasoningstructured actionsmultilingual interaction
- Known Limitations
- A running model server alone does not execute a closed-loop agent.Live-environment performance varies with app versions, task constraints, and execution scaffold.
- Notable Use Cases
- mobile agentsbrowser assistantsGUI-model research