Kolibri 1
Kolibri 1 is Aleph Alpha’s open-weight German-English reasoning and tool-use model, trained from scratch with a German-focused tokenizer and data mixture. Its hybrid-attention mixture-of-experts Transformer has 78B total parameters and approximately 3.46B active per token. The native 262,144-token context can extrapolate to 1,048,576 tokens, but the publisher recommends remaining within the native window for complex tasks and serving efficiency. Apache-licensed weights support private and sovereign deployments. Sparse activation reduces computation without eliminating the memory needed for the full weights; the published knowledge cutoff is June 18, 2026 for both English and German.
2026-10-03
78B total, approximately 3.46B active
Hybrid-attention Mixture-of-Experts Transformer
Apache-2.0
Specifications
- Parameters
- 78B total, approximately 3.46B active
- Architecture
- Hybrid-attention Mixture-of-Experts Transformer
- License
- Apache-2.0
- Context Window
- 1,048,576 tokens
- Training Data Cutoff
- 2026-06-18
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- Kolibri
- API Access
- Not Available
- Chat Interface
- Not Available
- Multilingual Support
- Yes
- Variants
- Kolibri-1FP8 serving
- Hardware Support
- CUDANVIDIA A100/H100/H200/B200/B300
Capabilities & Limitations
- Capabilities
- reasoningGerman and Englishtool callinglong context
- Known Limitations
- Native training context is 262,144; larger windows use extrapolationFull weights require substantially more memory than the active parameter countTraining cutoff limits implicit knowledge
- Notable Use Cases
- German enterprise assistantsprivate reasoning agentssovereign deployments
- Function Calling Support
- Yes
- Tool Use Support
- Yes