Xiaomi logo

MiMo V2.5

XiaomiOpen WeightsPending Human Review

MiMo V2.5 is Xiaomi’s full-modal language model with text, image, audio, and video understanding. Its sparse backbone has approximately 310B parameters and activates 15B per token, supporting a 1M-token context. It targets general agents, cross-modal reasoning, and tool-driven workflows. Public API testing preceded the April 27 weight publication. Both base and post-trained weights are now MIT-licensed; quantized serving and draft models are grouped rather than treated as new identities.

2026-04-23
310B total, 15B active
Hybrid-attention Mixture of Experts
MIT

Specifications

Parameters
310B total, 15B active
Architecture
Hybrid-attention Mixture of Experts
License
MIT
Context Window
1,048,576 tokens
Type
multimodal
Modalities
textimageaudiovideo

Benchmark Scores

Advanced Specifications

Model Family
MiMo
API Access
Available
Chat Interface
Available

Capabilities & Limitations

Capabilities
text generationreasoningmultimodal understanding
Known Limitations
Provider-reported evaluations need application-specific validationSparse or long-context serving can require substantial memoryCurrent information requires external retrieval
Notable Use Cases
private model deploymentdocument and code workflowslanguage-model research
Tool Use Support
Yes

Related Models