Ming-Flash-Omni 2.0
Ming-Flash-Omni2.0 is Ant Group’s language-capable any-to-any multimodal model, combining text, image, audio, and video understanding with text, image, and audio generation. Its Ling2.0-based MoE backbone has 100B total and 6B active parameters, with modality-specific encoders and generation components. The February 11 official release improves visual encyclopedic knowledge, natural speech, and image synthesis and editing. MIT weights and inference code are downloadable. An earlier Flash-Omni preview is grouped as a historical variant, while distinct Lite models have different sizes and specifications.
2026-02-11
100B total, 6B active
Ling2.0 Mixture of Experts backbone with multimodal encoders and generation modules
MIT
Specifications
- Parameters
- 100B total, 6B active
- Architecture
- Ling2.0 Mixture of Experts backbone with multimodal encoders and generation modules
- License
- MIT
- Type
- multimodal
- Modalities
- textimageaudiovideo
Benchmark Scores
Advanced Specifications
- Model Family
- Ming
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- Ming-Flash-Omni Preview (2025-10-27)
Capabilities & Limitations
- Capabilities
- multimodal reasoningvisual understandingspeech generationimage generationimage editing
- Known Limitations
- Different inference functions and modules handle understanding and generation.Generation and understanding quality depend on modality-specific input conditions.
- Notable Use Cases
- omnimodal assistantsaudio-video analysismultimodal research