MiMo V2.5
MiMo V2.5 is Xiaomi’s full-modal language model with text, image, audio, and video understanding. Its sparse backbone has approximately 310B parameters and activates 15B per token, supporting a 1M-token context. It targets general agents, cross-modal reasoning, and tool-driven workflows. Public API testing preceded the April 27 weight publication. Both base and post-trained weights are now MIT-licensed; quantized serving and draft models are grouped rather than treated as new identities.
2026-04-23
310B total, 15B active
Hybrid-attention Mixture of Experts
MIT
Specifications
- Parameters
- 310B total, 15B active
- Architecture
- Hybrid-attention Mixture of Experts
- License
- MIT
- Context Window
- 1,048,576 tokens
- Type
- multimodal
- Modalities
- textimageaudiovideo
Benchmark Scores
Advanced Specifications
- Model Family
- MiMo
- API Access
- Available
- Chat Interface
- Available
Capabilities & Limitations
- Capabilities
- text generationreasoningmultimodal understanding
- Known Limitations
- Provider-reported evaluations need application-specific validationSparse or long-context serving can require substantial memoryCurrent information requires external retrieval
- Notable Use Cases
- private model deploymentdocument and code workflowslanguage-model research
- Tool Use Support
- Yes