Inkling-Small
Matches its larger sibling on many benchmarks at a quarter of the size, with lower cost and latency.
Inkling-Small is a July 30, 2026 open-weights multimodal autoregressive transformer model release from Thinking Machines Lab. It is a sparse Mixture-of-Experts model with 276B total parameters and 12B active parameters, accepts text, image, and audio inputs, generates text outputs, and supports up to a 1M-token context window.
Pricing
Serverless inference beta pricing per 1M tokens for Inkling-Small on Tinker; page notes beta is available for Inkling and Inkling-Small only and production use requires joining a waitlist.
Thinking Machines' official Tinker Models & Pricing page lists Serverless Inference (Beta) pricing for Inkling-Small and states all prices are per million tokens. The Inkling-Small row shows Prefill (Input) $0.30, cached input $0.06, and Sample (Output) $1.20 for the Tinker ID thinkingmachines/Inkling-Small:peft:262144:sampling-nvfp4.