MiniCPM-SALA
MiniCPM-SALA is OpenBMB’s 9.48B language model combining sparse and linear attention for million-token contexts. One quarter of layers use InfLLM-v2 sparse attention and three quarters use Lightning Attention, with Hybrid Positional Embedding supporting length generalization beyond short training sequences. Apache-licensed weights and inference examples target long-context local language tasks with reduced cache overhead. The official model card documents deployment at up to 1M tokens, extending beyond its static 512K configuration value. Speedup claims depend on hardware, sequence length, and implementation rather than applying uniformly to every prompt.
2026-02-11
9.48B
Hybrid sparse/linear-attention language model
Apache-2.0
Specifications
- Parameters
- 9.48B
- Architecture
- Hybrid sparse/linear-attention language model
- License
- Apache-2.0
- Context Window
- 1,048,576 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- MiniCPM
- API Access
- Not Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- long contextefficient inferencereasoningcoding
- Known Limitations
- Million-token contexts require compatible positional and inference settingsAttention acceleration varies with sequence length and hardwareLong context does not guarantee accurate retrieval from every position
- Notable Use Cases
- large-document analysislocal long-context researchmemory-constrained language serving