InclusionAI (Ant Group) logo

Ming-UniVision

InclusionAI (Ant Group)Open WeightsPending Human Review

Ming-UniVision is a unified vision-language model that answers visual questions, describes images, generates images, and performs iterative image editing. A continuous MingTok vision tokenizer connects image understanding and generation with a single autoregressive next-token-prediction framework. The released 16B-A3B MoE checkpoint supports multi-round conversations that alternate between questions and editing requests. Public Apache-2.0 weights and code enable research and local inference. The October technical report establishes publication; the precise first public weight-release day is not asserted.

2025-10
16B total, 3B active
Autoregressive MoE vision-language model with unified continuous vision tokenizer
Apache-2.0

Specifications

Parameters
16B total, 3B active
Architecture
Autoregressive MoE vision-language model with unified continuous vision tokenizer
License
Apache-2.0
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
Ming
API Access
Not Available
Chat Interface
Not Available

Capabilities & Limitations

Capabilities
visual question answeringimage descriptionimage generationmulti-round image editing
Known Limitations
Requires the provided continuous-vision inference implementation.Visual generation and understanding need task-specific evaluation.
Notable Use Cases
image-grounded assistantsjoint understanding/generation researchinteractive image editing

Related Models