InclusionAI (Ant Group) logo

LLaDA2.1

InclusionAI (Ant Group)Open WeightsPending Human Review

LLaDA2.1 is Ant Group’s MoE diffusion-language model family, generating text through block-parallel refinement rather than exclusively left-to-right decoding. Token editing adds error-correcting generation and Speed/Quality modes, balancing parallel decoding speed against answer quality in 100B Flash and 16B Mini checkpoints. Public Apache-2.0 weights and custom inference code support research and self-hosting. Performance and latency depend on diffusion settings and the serving implementation.

2026-02-09
100B Flash; 16B Mini (non-embedding counts)
Mixture of Experts diffusion language model
Apache-2.0

Specifications

Parameters
100B Flash; 16B Mini (non-embedding counts)
Architecture
Mixture of Experts diffusion language model
License
Apache-2.0
Context Window
32,768 tokens
Type
text
Modalities
text

Benchmark Scores

Advanced Specifications

Model Family
LLaDA
API Access
Not Available
Chat Interface
Not Available
Variants
LLaDA2.1-flashLLaDA2.1-mini

Capabilities & Limitations

Capabilities
diffusion text generationcodinginstruction followingtoken editing
Known Limitations
Requires model-specific diffusion inference settings and runtime support.Speed and quality vary with block length, editing thresholds, and decoding mode.
Notable Use Cases
diffusion-language researchparallel text generationlocal coding assistants

Related Models