LLaDA2.0
LLaDA2.0 is Ant Group’s MoE diffusion-language model family, generating text through block-parallel refinement rather than exclusively left-to-right decoding. The release scales diffusion language modeling to a 100B Flash tier and a 16B Mini tier, with instruction-following and code-generation checkpoints and confidence-aware parallel variants. Public Apache-2.0 weights and custom inference code support research and self-hosting. Performance and latency depend on diffusion settings and the serving implementation.
2025-11
100B Flash; 16B Mini (non-embedding counts)
Mixture of Experts diffusion language model
Apache-2.0
Specifications
- Parameters
- 100B Flash; 16B Mini (non-embedding counts)
- Architecture
- Mixture of Experts diffusion language model
- License
- Apache-2.0
- Context Window
- 32,768 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- LLaDA
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- LLaDA2.0-flashLLaDA2.0-miniFlash/Mini CAP variants
Capabilities & Limitations
- Capabilities
- diffusion text generationcodinginstruction following
- Known Limitations
- Requires model-specific diffusion inference settings and runtime support.Speed and quality vary with block length, editing thresholds, and decoding mode.
- Notable Use Cases
- diffusion-language researchparallel text generationlocal coding assistants