OpenBMB logo

MiniCPM4

OpenBMBOpen WeightsPending Human Review

MiniCPM4 is OpenBMB’s efficient local language-model generation, with 8B and 0.5B checkpoints. It combines training-data improvements with sparse-attention and deployment optimizations intended for devices and constrained servers. Apache-licensed weights, code, and a technical report support local research and customization. This family page also groups BitCPM4 and the later BitCPM-CANN low-bit training variants derived from MiniCPM backbones. CANN checkpoints demonstrate ternary quantization-aware training on Huawei Ascend; their released pseudo-quantized floating-point storage does not itself provide a packed low-memory inference representation. The separately named MiniCPM4.1 update has its own page.

2025-06-06
0.5B and 8.19B
Decoder-only Transformer with sparse-attention optimizations
Apache-2.0

Specifications

Parameters
0.5B and 8.19B
Architecture
Decoder-only Transformer with sparse-attention optimizations
License
Apache-2.0
Type
text
Modalities
text

Benchmark Scores

Advanced Specifications

Model Family
MiniCPM
API Access
Not Available
Chat Interface
Not Available
Variants
MiniCPM4-0.5BMiniCPM4-8BBitCPM4-0.5BBitCPM4-1BBitCPM-CANN-0.5BBitCPM-CANN-1BBitCPM-CANN-3BBitCPM-CANN-8B

Capabilities & Limitations

Capabilities
local language generationcodingefficient inferenceon-device use
Known Limitations
Size and variant behavior differLow-bit training variants require compatible inference implementations for memory savingsModel outputs still require factual and code validation
Notable Use Cases
offline assistantsresource-constrained language applicationslow-bit model research

Related Models