LLaDA2.0-Uni
LLaDA2.0-Uni unifies visual understanding, image generation, and instruction-based editing within a diffusion language model. Built from LLaDA2.0-Mini, its MoE backbone predicts masked semantic tokens, while SigLIP-VQ encodes images and a diffusion decoder reconstructs visual outputs. It supports visual question answering, image captioning, document understanding, and preliminary interleaved reasoning and generation. Apache-2.0 weights and code were released April 23, with FP8 and additional serving integrations arriving later. Its language-based visual understanding places it within the multimodal language-model scope.
2026-04-23
Diffusion MoE language backbone with semantic image tokenizer and image diffusion decoder
Apache-2.0
Specifications
- Architecture
- Diffusion MoE language backbone with semantic image tokenizer and image diffusion decoder
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- LLaDA
- Finetuned From
- LLaDA2.0-mini
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- FP8 distribution
Capabilities & Limitations
- Capabilities
- visual question answeringimage captioningimage generationimage editingmultimodal reasoning
- Known Limitations
- Interleaved generation support is preliminary.Specialized visual inference code and dependencies are required.
- Notable Use Cases
- multimodal researchimage-grounded assistantscombined visual understanding and generation