OpenBMB logo

MiniCPM-V 4.5

OpenBMBOpen WeightsPending Human Review

MiniCPM-V 4.5 is OpenBMB’s open-weight vision-language model based on Qwen3-8B and SigLIP2-400M. It supports text, single or multiple images, and video understanding with language output. High-resolution image processing and compressed video tokens target document questions, visual reasoning, and longer visual sequences on local devices. It offers thinking and instruction behavior alongside multilingual capabilities spanning more than 30 languages. Apache-licensed weights, inference examples, and a demo are available. Reported OCR and document benchmarks reflect particular test conditions and do not make extracted information automatically reliable.

2025-08-26
8B
Vision-language Transformer with compressed visual tokens
Apache-2.0

Specifications

Parameters
8B
Architecture
Vision-language Transformer with compressed visual tokens
License
Apache-2.0
Type
multimodal
Modalities
textimagevideo

Benchmark Scores

Advanced Specifications

Model Family
MiniCPM
Finetuned From
Qwen3-8B
API Access
Not Available
Chat Interface
Available
Multilingual Support
Yes

Capabilities & Limitations

Capabilities
visual question answeringvideo understandingdocument analysisreasoningmultilingual
Known Limitations
High-resolution or multi-frame inputs increase runtime workVisual and document interpretations can be incorrectLong visual sequences depend on sampling and token-compression settings
Notable Use Cases
local visual assistantsdocument question answeringvideo reasoning research

Related Models