OpenBMB logo

MiniCPM-SALA

OpenBMBOpen WeightsPending Human Review

MiniCPM-SALA is OpenBMB’s 9.48B language model combining sparse and linear attention for million-token contexts. One quarter of layers use InfLLM-v2 sparse attention and three quarters use Lightning Attention, with Hybrid Positional Embedding supporting length generalization beyond short training sequences. Apache-licensed weights and inference examples target long-context local language tasks with reduced cache overhead. The official model card documents deployment at up to 1M tokens, extending beyond its static 512K configuration value. Speedup claims depend on hardware, sequence length, and implementation rather than applying uniformly to every prompt.

2026-02-11
9.48B
Hybrid sparse/linear-attention language model
Apache-2.0

Specifications

Parameters
9.48B
Architecture
Hybrid sparse/linear-attention language model
License
Apache-2.0
Context Window
1,048,576 tokens
Type
text
Modalities
text

Benchmark Scores

Advanced Specifications

Model Family
MiniCPM
API Access
Not Available
Chat Interface
Not Available

Capabilities & Limitations

Capabilities
long contextefficient inferencereasoningcoding
Known Limitations
Million-token contexts require compatible positional and inference settingsAttention acceleration varies with sequence length and hardwareLong context does not guarantee accurate retrieval from every position
Notable Use Cases
large-document analysislocal long-context researchmemory-constrained language serving

Related Models