MiMo V2 Omni
MiMo V2 Omni is Xiaomi’s native multimodal model for visual, audio, video, and text understanding. It combines modality encoders with a shared language backbone so an agent can ground its responses in screenshots, recordings, and other inputs. Hosted inference supports reasoning and structured tool use. It is a distinct multimodal tier of the V2 generation, with parameter counts left unspecified because the launch materials do not disclose a reliable total. Its availability does not imply downloadable weights.
2026-03-18
Multimodal Transformer
Proprietary
Specifications
- Architecture
- Multimodal Transformer
- License
- Proprietary
- Context Window
- 262,144 tokens
- Type
- multimodal
- Modalities
- textimageaudiovideo
Benchmark Scores
Advanced Specifications
- Model Family
- MiMo
- API Access
- Available
- Chat Interface
- Available
Capabilities & Limitations
- Capabilities
- text generationreasoningmultimodal understanding
- Known Limitations
- Provider-reported evaluations need application-specific validationGenerated interpretations can be incorrectCustom license conditions apply
- Notable Use Cases
- hosted assistantsdocument and code workflowslanguage-model research
- Tool Use Support
- Yes