InclusionAI (Ant Group) logo

UI-Venus-2

InclusionAI (Ant Group)Open WeightsPending Human Review

UI-Venus-2 is a distinct vision-language model specialization for GUI interaction, combining screenshot understanding with natural-language task reasoning and structured actions. The second generation adds multilingual mobile, browser, and desktop environments, keypoint-grounded verification, and reflection-based recovery. The released 9B checkpoint is initialized from Qwen3.5-9B; a 27B tier is described in research evaluations, but its weights are not verified. The current model card explicitly leaves its weight license unresolved. Local inference requires an external action parser and execution framework; the model itself emits language reasoning and actions.

2026-08-27
9B released weights; 27B published research tier
Transformer vision-language model based on Qwen3.5-9B
Undisclosed: model-weight license pending confirmation

Specifications

Parameters
9B released weights; 27B published research tier
Architecture
Transformer vision-language model based on Qwen3.5-9B
License
Undisclosed: model-weight license pending confirmation
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
UI-Venus
Finetuned From
Qwen3.5-9B
API Access
Not Available
Chat Interface
Not Available

Capabilities & Limitations

Capabilities
GUI groundingtask reasoningstructured actionsmultilingual interaction
Known Limitations
A running model server alone does not execute a closed-loop agent.Live-environment performance varies with app versions, task constraints, and execution scaffold.The weight license is explicitly pending confirmation.
Notable Use Cases
mobile agentsbrowser assistantsGUI-model research

Related Models