InclusionAI (Ant Group) logo

LLaDA2.0-Uni

InclusionAI (Ant Group)Open WeightsPending Human Review

LLaDA2.0-Uni unifies visual understanding, image generation, and instruction-based editing within a diffusion language model. Built from LLaDA2.0-Mini, its MoE backbone predicts masked semantic tokens, while SigLIP-VQ encodes images and a diffusion decoder reconstructs visual outputs. It supports visual question answering, image captioning, document understanding, and preliminary interleaved reasoning and generation. Apache-2.0 weights and code were released April 23, with FP8 and additional serving integrations arriving later. Its language-based visual understanding places it within the multimodal language-model scope.

2026-04-23
Diffusion MoE language backbone with semantic image tokenizer and image diffusion decoder
Apache-2.0

Specifications

Architecture
Diffusion MoE language backbone with semantic image tokenizer and image diffusion decoder
License
Apache-2.0
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
LLaDA
Finetuned From
LLaDA2.0-mini
API Access
Not Available
Chat Interface
Not Available
Variants
FP8 distribution

Capabilities & Limitations

Capabilities
visual question answeringimage captioningimage generationimage editingmultimodal reasoning
Known Limitations
Interleaved generation support is preliminary.Specialized visual inference code and dependencies are required.
Notable Use Cases
multimodal researchimage-grounded assistantscombined visual understanding and generation

Related Models