MiniCPM4.1
MiniCPM4.1 is OpenBMB’s 8B language model with trainable sparse attention and hybrid reasoning. InfLLM-v2 enables efficient attention to long text, while fusion-thinking post-training supports both deliberative and direct answers. The native context is 64K; the model card separately validates 128K with LongRoPE scaling, which is recorded as an extension rather than the default. Apache-licensed weights and local deployment code are available. Its efficiency focus suits private assistants and coding applications on devices or modest servers. Sparse attention settings, reasoning mode, and context scaling affect both quality and runtime behavior.
2025-09-05
8B
Sparse-attention decoder-only Transformer
Apache-2.0
Specifications
- Parameters
- 8B
- Architecture
- Sparse-attention decoder-only Transformer
- License
- Apache-2.0
- Context Window
- 65,536 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- MiniCPM
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- MiniCPM4.1-8BLongRoPE 128K extension
Capabilities & Limitations
- Capabilities
- hybrid reasoningcodinglong contextefficient local inference
- Known Limitations
- 128K use requires the documented LongRoPE extensionSparse-attention settings affect correctness and efficiencyReasoning traces increase token and latency costs
- Notable Use Cases
- private assistantslocal coding supportlong-text reasoning research