Inkling
Thinking Machines Lab

Inkling is Thinking Machines Lab’s open-weight multimodal mixture-of-experts model, trained from scratch for native reasoning over text, images, and audio. Variable thinking effort allows applications to choose a latency and reasoning tradeoff. The checkpoint supports up to 1M tokens, while hosted product limits can be lower. Apache-licensed weights and Tinker access enable customization and deployment for coding, conversational assistants, and tool-using agents. Its pretraining includes 45T tokens spanning text, images, audio, and video; video training does not establish a supported direct video-input endpoint.
Typemultimodal
Parameters975B total, 41B active