Ming-UniVision
Ming-UniVision is a unified vision-language model that answers visual questions, describes images, generates images, and performs iterative image editing. A continuous MingTok vision tokenizer connects image understanding and generation with a single autoregressive next-token-prediction framework. The released 16B-A3B MoE checkpoint supports multi-round conversations that alternate between questions and editing requests. Public Apache-2.0 weights and code enable research and local inference. The October technical report establishes publication; the precise first public weight-release day is not asserted.
2025-10
16B total, 3B active
Autoregressive MoE vision-language model with unified continuous vision tokenizer
Apache-2.0
Specifications
- Parameters
- 16B total, 3B active
- Architecture
- Autoregressive MoE vision-language model with unified continuous vision tokenizer
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- Ming
- API Access
- Not Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- visual question answeringimage descriptionimage generationmulti-round image editing
- Known Limitations
- Requires the provided continuous-vision inference implementation.Visual generation and understanding need task-specific evaluation.
- Notable Use Cases
- image-grounded assistantsjoint understanding/generation researchinteractive image editing