MiniCPM-V 4.5
MiniCPM-V 4.5 is OpenBMB’s open-weight vision-language model based on Qwen3-8B and SigLIP2-400M. It supports text, single or multiple images, and video understanding with language output. High-resolution image processing and compressed video tokens target document questions, visual reasoning, and longer visual sequences on local devices. It offers thinking and instruction behavior alongside multilingual capabilities spanning more than 30 languages. Apache-licensed weights, inference examples, and a demo are available. Reported OCR and document benchmarks reflect particular test conditions and do not make extracted information automatically reliable.
2025-08-26
8B
Vision-language Transformer with compressed visual tokens
Apache-2.0
Specifications
- Parameters
- 8B
- Architecture
- Vision-language Transformer with compressed visual tokens
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimagevideo
Benchmark Scores
Advanced Specifications
- Model Family
- MiniCPM
- Finetuned From
- Qwen3-8B
- API Access
- Not Available
- Chat Interface
- Available
- Multilingual Support
- Yes
Capabilities & Limitations
- Capabilities
- visual question answeringvideo understandingdocument analysisreasoningmultilingual
- Known Limitations
- High-resolution or multi-frame inputs increase runtime workVisual and document interpretations can be incorrectLong visual sequences depend on sampling and token-compression settings
- Notable Use Cases
- local visual assistantsdocument question answeringvideo reasoning research