LLaDA2.1
LLaDA2.1 is Ant Group’s MoE diffusion-language model family, generating text through block-parallel refinement rather than exclusively left-to-right decoding. Token editing adds error-correcting generation and Speed/Quality modes, balancing parallel decoding speed against answer quality in 100B Flash and 16B Mini checkpoints. Public Apache-2.0 weights and custom inference code support research and self-hosting. Performance and latency depend on diffusion settings and the serving implementation.
2026-02-09
100B Flash; 16B Mini (non-embedding counts)
Mixture of Experts diffusion language model
Apache-2.0
Specifications
- Parameters
- 100B Flash; 16B Mini (non-embedding counts)
- Architecture
- Mixture of Experts diffusion language model
- License
- Apache-2.0
- Context Window
- 32,768 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- LLaDA
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- LLaDA2.1-flashLLaDA2.1-mini
Capabilities & Limitations
- Capabilities
- diffusion text generationcodinginstruction followingtoken editing
- Known Limitations
- Requires model-specific diffusion inference settings and runtime support.Speed and quality vary with block length, editing thresholds, and decoding mode.
- Notable Use Cases
- diffusion-language researchparallel text generationlocal coding assistants