Inkling Small
Inkling Small is Thinking Machines Lab’s open-weight multimodal mixture-of-experts model, trained from scratch for native reasoning over text, images, and audio. Variable thinking effort allows applications to choose a latency and reasoning tradeoff. The checkpoint supports up to 1M tokens, while hosted product limits can be lower. Apache-licensed weights and Tinker access enable customization and deployment for coding, conversational assistants, and tool-using agents. The July 15 announcement previewed the Small tier; its full weight release and model card are dated July 30.
2026-07-30
276B total, 12B active
Mixture-of-Experts Transformer
Apache-2.0
Specifications
- Parameters
- 276B total, 12B active
- Architecture
- Mixture-of-Experts Transformer
- License
- Apache-2.0
- Context Window
- 1,048,576 tokens
- Type
- multimodal
- Modalities
- textimageaudio
Benchmark Scores
Advanced Specifications
- Model Family
- Inkling
- API Access
- Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- multimodal reasoningcodingtool useadjustable reasoninglong context
- Known Limitations
- Full weights require substantial distributed-serving memoryHosted Tinker context limits differ from checkpoint capacityModel-card evaluations are provider-reported
- Notable Use Cases
- customizable assistantsmultimodal agentsprivate model research
- Tool Use Support
- Yes