MiniCPM-V 4.6
MiniCPM-V 4.6 is OpenBMB’s compact 1.3B vision-language model for efficient device deployment. It combines Qwen3.5-0.8B with SigLIP2-400M and mixed 4x/16x visual-token compression, accepting text, images, and video and producing language answers. The release focuses on reducing visual encoding work while preserving multi-image and video understanding. Apache-licensed weights, ordinary and Thinking checkpoints, a demo, and a later May 17 API are available. This page does not equate the backbone configuration maximum with validated long-context visual use. Throughput comparisons are provider measurements tied to particular input and serving conditions.
2026-05-11
1.3B
Hybrid-attention vision-language model with visual-token compression
Apache-2.0
Specifications
- Parameters
- 1.3B
- Architecture
- Hybrid-attention vision-language model with visual-token compression
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimagevideo
Benchmark Scores
Advanced Specifications
- Model Family
- MiniCPM
- Finetuned From
- Qwen3.5-0.8B
- API Access
- Available
- Chat Interface
- Available
- Variants
- MiniCPM-V-4.6MiniCPM-V-4.6-ThinkingGGUFAWQBNB
Capabilities & Limitations
- Capabilities
- visual question answeringvideo understandingon-device inferencedocument analysis
- Known Limitations
- Compact models can fail on difficult reasoning or small visual detailsCompression changes the fidelity and cost of visual inputsProvider throughput claims require matching serving conditions
- Notable Use Cases
- device visual assistantslocal document questionsefficient video-analysis research