LLaDA-UI
LLaDA-UI brings block-wise diffusion generation to a vision-language GUI model. The approximately 16.7B MoE model preserves screenshot aspect ratios and visual detail, reasons about natural-language tasks, and produces grounded coordinates or tagged reasoning and structured actions for mobile, desktop, and web interfaces. The September 9 technical report precedes the September 15 weight publication, independently of the repository creation timestamp. Public weights and serving examples are available; the exact weight license is not declared in the current card and is not inferred from neighboring LLaDA models.
2026-09-09
Approximately 16.7B
Block-wise diffusion MoE vision-language model
Undisclosed
Specifications
- Parameters
- Approximately 16.7B
- Architecture
- Block-wise diffusion MoE vision-language model
- License
- Undisclosed
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- LLaDA
- Finetuned From
- LLaDA2.0-mini-base
- API Access
- Not Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- GUI reasoningvisual groundingstructured action generationblock-wise diffusion
- Known Limitations
- Requires model-specific diffusion serving and an external action executor.The current card does not establish a model-weight license.
- Notable Use Cases
- GUI-model researchmobile assistantsdesktop agent prototypes