EMO
EMO is Ai2’s open mixture-of-experts language model trained so reusable domain modules emerge without predefined semantic labels. Its 128 routed experts activate eight per token, giving 14B total and about 1B active parameters. A document-level routing objective encourages coherent expert groups that can be selected for particular tasks. The main checkpoint is trained on one trillion tokens with additional annealing. Ai2 releases weights, matched baselines, and training code for modularity research. Results for reduced expert subsets depend on the selection method and evaluation task.
2026-05-08
14B total, ~1B active
Sparse mixture-of-experts Transformer with document-level expert-pool training
Apache-2.0
Specifications
- Parameters
- 14B total, ~1B active
- Architecture
- Sparse mixture-of-experts Transformer with document-level expert-pool training
- License
- Apache-2.0
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- EMO
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- Main 1T checkpoint130B-token ablation
Capabilities & Limitations
- Capabilities
- text generationmodular expert selection
- Known Limitations
- Research completion modelSelecting a small expert subset can reduce performance
- Notable Use Cases
- modular language-model researchefficient specialist adaptation