UI-Venus-2
UI-Venus-2 is a distinct vision-language model specialization for GUI interaction, combining screenshot understanding with natural-language task reasoning and structured actions. The second generation adds multilingual mobile, browser, and desktop environments, keypoint-grounded verification, and reflection-based recovery. The released 9B checkpoint is initialized from Qwen3.5-9B; a 27B tier is described in research evaluations, but its weights are not verified. The current model card explicitly leaves its weight license unresolved. Local inference requires an external action parser and execution framework; the model itself emits language reasoning and actions.
2026-08-27
9B released weights; 27B published research tier
Transformer vision-language model based on Qwen3.5-9B
Undisclosed: model-weight license pending confirmation
Specifications
- Parameters
- 9B released weights; 27B published research tier
- Architecture
- Transformer vision-language model based on Qwen3.5-9B
- License
- Undisclosed: model-weight license pending confirmation
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- UI-Venus
- Finetuned From
- Qwen3.5-9B
- API Access
- Not Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- GUI groundingtask reasoningstructured actionsmultilingual interaction
- Known Limitations
- A running model server alone does not execute a closed-loop agent.Live-environment performance varies with app versions, task constraints, and execution scaffold.The weight license is explicitly pending confirmation.
- Notable Use Cases
- mobile agentsbrowser assistantsGUI-model research