InclusionAI (Ant Group) logo

LLaDA-UI

InclusionAI (Ant Group)Open WeightsPending Human Review

LLaDA-UI brings block-wise diffusion generation to a vision-language GUI model. The approximately 16.7B MoE model preserves screenshot aspect ratios and visual detail, reasons about natural-language tasks, and produces grounded coordinates or tagged reasoning and structured actions for mobile, desktop, and web interfaces. The September 9 technical report precedes the September 15 weight publication, independently of the repository creation timestamp. Public weights and serving examples are available; the exact weight license is not declared in the current card and is not inferred from neighboring LLaDA models.

2026-09-09
Approximately 16.7B
Block-wise diffusion MoE vision-language model
Undisclosed

Specifications

Parameters
Approximately 16.7B
Architecture
Block-wise diffusion MoE vision-language model
License
Undisclosed
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
LLaDA
Finetuned From
LLaDA2.0-mini-base
API Access
Not Available
Chat Interface
Not Available

Capabilities & Limitations

Capabilities
GUI reasoningvisual groundingstructured action generationblock-wise diffusion
Known Limitations
Requires model-specific diffusion serving and an external action executor.The current card does not establish a model-weight license.
Notable Use Cases
GUI-model researchmobile assistantsdesktop agent prototypes

Related Models