Ming-Lite-Omni
Ming-Lite-Omni is Ant Group’s compact omnimodal language-model family, combining image, text, video, and speech understanding with text, speech, and image generation. The public preview was released May 4, followed by the May 28 official checkpoint and July 15 v1.5 update. The v1.5 model has 20.3B total parameters with 3B active in its MoE backbone, adding stronger video comprehension, refined image editing, and improved speech synthesis. MIT weights and model-specific inference code enable self-hosting. This page groups the original Lite sequence; the substantially larger Flash-Omni2.0 tier has a separate page.
2025-05-04
20.3B total; 3B active in MoE backbone (v1.5)
MoE language backbone with multimodal encoders and generation components
MIT
Specifications
- Parameters
- 20.3B total; 3B active in MoE backbone (v1.5)
- Architecture
- MoE language backbone with multimodal encoders and generation components
- License
- MIT
- Type
- multimodal
- Modalities
- textimageaudiovideo
Benchmark Scores
Advanced Specifications
- Model Family
- Ming
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- Ming-Lite-Omni Preview (2025-05-04)Ming-Lite-Omni v1 (2025-05-28)Ming-Lite-Omni v1.5 (2025-07-15)
Capabilities & Limitations
- Capabilities
- multimodal reasoningvisual question answeringspeech generationimage generationimage editing
- Known Limitations
- Version-specific inference modules are required.Sizes and capabilities vary between preview, v1, and v1.5 checkpoints.
- Notable Use Cases
- local multimodal assistantsaudio-video analysismultimodal research