Xiaomi logo

MiMo V2.6 Pro

XiaomiOpen WeightsPending Human Review

MiMo V2.6 Pro extends Xiaomi’s flagship into native multimodal agent reasoning over text, images, audio, and video. It retains a 1.02T sparse backbone with 42B active parameters and a 1M-token context. Large-scale reinforcement learning uses verifiable coding, knowledge-work, visual, and security environments. API access and MIT-licensed RL checkpoints were released together. Later MOPD refinements and UltraSpeed serving are grouped as variants; benchmark results depend on the chosen checkpoint and agent harness.

2026-09-22
1.02T total, 42B active
Hybrid-attention Multimodal Mixture of Experts
MIT

Specifications

Parameters
1.02T total, 42B active
Architecture
Hybrid-attention Multimodal Mixture of Experts
License
MIT
Context Window
1,048,576 tokens
Type
multimodal
Modalities
textimageaudiovideo

Benchmark Scores

Advanced Specifications

Model Family
MiMo
API Access
Available
Chat Interface
Available

Capabilities & Limitations

Capabilities
text generationreasoningmultimodal understanding
Known Limitations
Provider-reported evaluations need application-specific validationSparse or long-context serving can require substantial memoryCurrent information requires external retrieval
Notable Use Cases
private model deploymentdocument and code workflowslanguage-model research
Tool Use Support
Yes

Related Models