Granite SWASH
Granite SWASH is IBM Research’s open architectural preview for efficient language modeling. The family explores sliding-window attention at every layer, learnable attention sinks, and hybrid transformer designs, with dense 2B and sparse 3B checkpoints. The sparse model activates approximately 600M parameters. These are general-purpose English base models for research and customization, rather than instruction-ready assistants. Apache-licensed weights allow study of the architecture and downstream training; chat and tool behavior should not be inferred from the base release.
2026-07-07
2B dense; 3B total, 600M active
Sliding-window Transformer with learned attention sinks
Apache-2.0
Specifications
- Parameters
- 2B dense; 3B total, 600M active
- Architecture
- Sliding-window Transformer with learned attention sinks
- License
- Apache-2.0
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- Granite
- API Access
- Not Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- text generationbase-model research
- Known Limitations
- Provider-reported evaluations need application-specific validationGenerated interpretations can be incorrectCurrent information requires external retrieval
- Notable Use Cases
- private model deploymentdocument and code workflowslanguage-model research