OpenBMB logo

MiniCPM4.1

OpenBMBOpen WeightsPending Human Review

MiniCPM4.1 is OpenBMB’s 8B language model with trainable sparse attention and hybrid reasoning. InfLLM-v2 enables efficient attention to long text, while fusion-thinking post-training supports both deliberative and direct answers. The native context is 64K; the model card separately validates 128K with LongRoPE scaling, which is recorded as an extension rather than the default. Apache-licensed weights and local deployment code are available. Its efficiency focus suits private assistants and coding applications on devices or modest servers. Sparse attention settings, reasoning mode, and context scaling affect both quality and runtime behavior.

2025-09-05
8B
Sparse-attention decoder-only Transformer
Apache-2.0

Specifications

Parameters
8B
Architecture
Sparse-attention decoder-only Transformer
License
Apache-2.0
Context Window
65,536 tokens
Type
text
Modalities
text

Benchmark Scores

Advanced Specifications

Model Family
MiniCPM
API Access
Not Available
Chat Interface
Not Available
Variants
MiniCPM4.1-8BLongRoPE 128K extension

Capabilities & Limitations

Capabilities
hybrid reasoningcodinglong contextefficient local inference
Known Limitations
128K use requires the documented LongRoPE extensionSparse-attention settings affect correctness and efficiencyReasoning traces increase token and latency costs
Notable Use Cases
private assistantslocal coding supportlong-text reasoning research

Related Models