MiMo V2 Flash
MiMo V2 Flash is Xiaomi’s open-weight sparse language model for efficient reasoning, coding, and tool-using agents. It has 309B total parameters and 15B active per token, with hybrid attention and multi-token prediction improving serving efficiency. Downloadable base and post-trained weights use MIT licensing. A 256K context accommodates large codebases and documents. The active parameter count describes computation per token; deployment still needs storage for the entire sparse model.
2025-12-16
309B total, 15B active
Hybrid-attention Mixture of Experts
MIT
Specifications
- Parameters
- 309B total, 15B active
- Architecture
- Hybrid-attention Mixture of Experts
- License
- MIT
- Context Window
- 262,144 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- MiMo
- API Access
- Available
- Chat Interface
- Available
Capabilities & Limitations
- Capabilities
- text generationreasoning
- Known Limitations
- Provider-reported evaluations need application-specific validationSparse or long-context serving can require substantial memoryCurrent information requires external retrieval
- Notable Use Cases
- private model deploymentdocument and code workflowslanguage-model research
- Tool Use Support
- Yes