MiniCPM4
MiniCPM4 is OpenBMB’s efficient local language-model generation, with 8B and 0.5B checkpoints. It combines training-data improvements with sparse-attention and deployment optimizations intended for devices and constrained servers. Apache-licensed weights, code, and a technical report support local research and customization. This family page also groups BitCPM4 and the later BitCPM-CANN low-bit training variants derived from MiniCPM backbones. CANN checkpoints demonstrate ternary quantization-aware training on Huawei Ascend; their released pseudo-quantized floating-point storage does not itself provide a packed low-memory inference representation. The separately named MiniCPM4.1 update has its own page.
2025-06-06
0.5B and 8.19B
Decoder-only Transformer with sparse-attention optimizations
Apache-2.0
Specifications
- Parameters
- 0.5B and 8.19B
- Architecture
- Decoder-only Transformer with sparse-attention optimizations
- License
- Apache-2.0
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- MiniCPM
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- MiniCPM4-0.5BMiniCPM4-8BBitCPM4-0.5BBitCPM4-1BBitCPM-CANN-0.5BBitCPM-CANN-1BBitCPM-CANN-3BBitCPM-CANN-8B
Capabilities & Limitations
- Capabilities
- local language generationcodingefficient inferenceon-device use
- Known Limitations
- Size and variant behavior differLow-bit training variants require compatible inference implementations for memory savingsModel outputs still require factual and code validation
- Notable Use Cases
- offline assistantsresource-constrained language applicationslow-bit model research