Tülu 3
Tülu 3 is Ai2’s open post-training model family built on Meta Llama 3.1. It combines curated supervised instruction data, direct preference optimization, and reinforcement learning with verifiable rewards to improve chat, mathematical reasoning, instruction following, and tool use. The 8B and 70B releases were followed by a 405B tier and a 3.1 post-training refresh. Training data, scripts, and intermediate stages are public, but underlying Llama weight-license conditions still apply. This family page groups the routine size and post-training-stage checkpoints.
2024-11-21
8B, 70B, 405B
Llama-derived decoder-only Transformer
Llama 3.1 Community License; released training artifacts have separate licenses
Specifications
- Parameters
- 8B, 70B, 405B
- Architecture
- Llama-derived decoder-only Transformer
- License
- Llama 3.1 Community License; released training artifacts have separate licenses
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- Tülu
- Finetuned From
- Llama 3.1
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- 8B70B405B (January 2025)3.1 refresh
Capabilities & Limitations
- Capabilities
- instruction followingreasoningmathtool use
- Known Limitations
- Underlying Llama license appliesTraining stages and sizes have different behavior
- Notable Use Cases
- post-training researchself-hosted chatalignment experiments
- Tool Use Support
- Yes