LLaDA2.2
LLaDA2.2 is Ant Group’s MoE diffusion-language model family, generating text through block-parallel refinement rather than exclusively left-to-right decoding. Levenshtein editing introduces DELETE and INSERT tokens for structural correction, while block routing and agentic reinforcement learning support 128K tool use and multi-turn interaction. The 100B Flash checkpoint was followed by the 16B Mini release in September. Public Apache-2.0 weights and custom inference code support research and self-hosting. Performance and latency depend on diffusion settings and the serving implementation.
2026-07
100B Flash; 16B Mini (non-embedding counts)
Mixture of Experts diffusion language model
Apache-2.0
Specifications
- Parameters
- 100B Flash; 16B Mini (non-embedding counts)
- Architecture
- Mixture of Experts diffusion language model
- License
- Apache-2.0
- Context Window
- 131,072 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- LLaDA
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- LLaDA2.2-flashLLaDA2.2-mini (2026-09 release)
Capabilities & Limitations
- Capabilities
- diffusion text generationcodinginstruction followingtoken editingagentic tool use
- Known Limitations
- Requires model-specific diffusion inference settings and runtime support.Speed and quality vary with block length, editing thresholds, and decoding mode.
- Notable Use Cases
- diffusion-language researchparallel text generationlocal coding assistants
- Function Calling Support
- Yes
- Tool Use Support
- Yes