OpenBMB logo

MiniCPM-V 4.6

OpenBMBOpen WeightsPending Human Review

MiniCPM-V 4.6 is OpenBMB’s compact 1.3B vision-language model for efficient device deployment. It combines Qwen3.5-0.8B with SigLIP2-400M and mixed 4x/16x visual-token compression, accepting text, images, and video and producing language answers. The release focuses on reducing visual encoding work while preserving multi-image and video understanding. Apache-licensed weights, ordinary and Thinking checkpoints, a demo, and a later May 17 API are available. This page does not equate the backbone configuration maximum with validated long-context visual use. Throughput comparisons are provider measurements tied to particular input and serving conditions.

2026-05-11
1.3B
Hybrid-attention vision-language model with visual-token compression
Apache-2.0

Specifications

Parameters
1.3B
Architecture
Hybrid-attention vision-language model with visual-token compression
License
Apache-2.0
Type
multimodal
Modalities
textimagevideo

Benchmark Scores

Advanced Specifications

Model Family
MiniCPM
Finetuned From
Qwen3.5-0.8B
API Access
Available
Chat Interface
Available
Variants
MiniCPM-V-4.6MiniCPM-V-4.6-ThinkingGGUFAWQBNB

Capabilities & Limitations

Capabilities
visual question answeringvideo understandingon-device inferencedocument analysis
Known Limitations
Compact models can fail on difficult reasoning or small visual detailsCompression changes the fidelity and cost of visual inputsProvider throughput claims require matching serving conditions
Notable Use Cases
device visual assistantslocal document questionsefficient video-analysis research

Related Models