Xiaomi logo

MiMo V2 Omni

XiaomiProprietaryPending Human Review

MiMo V2 Omni is Xiaomi’s native multimodal model for visual, audio, video, and text understanding. It combines modality encoders with a shared language backbone so an agent can ground its responses in screenshots, recordings, and other inputs. Hosted inference supports reasoning and structured tool use. It is a distinct multimodal tier of the V2 generation, with parameter counts left unspecified because the launch materials do not disclose a reliable total. Its availability does not imply downloadable weights.

2026-03-18
Multimodal Transformer
Proprietary

Specifications

Architecture
Multimodal Transformer
License
Proprietary
Context Window
262,144 tokens
Type
multimodal
Modalities
textimageaudiovideo

Benchmark Scores

Advanced Specifications

Model Family
MiMo
API Access
Available
Chat Interface
Available

Capabilities & Limitations

Capabilities
text generationreasoningmultimodal understanding
Known Limitations
Provider-reported evaluations need application-specific validationGenerated interpretations can be incorrectCustom license conditions apply
Notable Use Cases
hosted assistantsdocument and code workflowslanguage-model research
Tool Use Support
Yes

Related Models