Inkling
Inkling is Thinking Machines Lab’s open-weight multimodal mixture-of-experts model, trained from scratch for native reasoning over text, images, and audio. Variable thinking effort allows applications to choose a latency and reasoning tradeoff. The checkpoint supports up to 1M tokens, while hosted product limits can be lower. Apache-licensed weights and Tinker access enable customization and deployment for coding, conversational assistants, and tool-using agents. Its pretraining includes 45T tokens spanning text, images, audio, and video; video training does not establish a supported direct video-input endpoint.
2026-07-15
975B total, 41B active
Mixture-of-Experts Transformer
Apache-2.0
Specifications
- Parameters
- 975B total, 41B active
- Architecture
- Mixture-of-Experts Transformer
- License
- Apache-2.0
- Context Window
- 1,048,576 tokens
- Type
- multimodal
- Modalities
- textimageaudio
Benchmark Scores
Advanced Specifications
- Model Family
- Inkling
- API Access
- Available
- Chat Interface
- Not Available
Capabilities & Limitations
- Capabilities
- multimodal reasoningcodingtool useadjustable reasoninglong context
- Known Limitations
- Full weights require substantial distributed-serving memoryHosted Tinker context limits differ from checkpoint capacityModel-card evaluations are provider-reported
- Notable Use Cases
- customizable assistantsmultimodal agentsprivate model research
- Tool Use Support
- Yes