MiMo V2.6 Pro
MiMo V2.6 Pro extends Xiaomi’s flagship into native multimodal agent reasoning over text, images, audio, and video. It retains a 1.02T sparse backbone with 42B active parameters and a 1M-token context. Large-scale reinforcement learning uses verifiable coding, knowledge-work, visual, and security environments. API access and MIT-licensed RL checkpoints were released together. Later MOPD refinements and UltraSpeed serving are grouped as variants; benchmark results depend on the chosen checkpoint and agent harness.
2026-09-22
1.02T total, 42B active
Hybrid-attention Multimodal Mixture of Experts
MIT
Specifications
- Parameters
- 1.02T total, 42B active
- Architecture
- Hybrid-attention Multimodal Mixture of Experts
- License
- MIT
- Context Window
- 1,048,576 tokens
- Type
- multimodal
- Modalities
- textimageaudiovideo
Benchmark Scores
Advanced Specifications
- Model Family
- MiMo
- API Access
- Available
- Chat Interface
- Available
Capabilities & Limitations
- Capabilities
- text generationreasoningmultimodal understanding
- Known Limitations
- Provider-reported evaluations need application-specific validationSparse or long-context serving can require substantial memoryCurrent information requires external retrieval
- Notable Use Cases
- private model deploymentdocument and code workflowslanguage-model research
- Tool Use Support
- Yes