Skip to content
AI Primer

Inkling-Small

Matches its larger sibling on many benchmarks at a quarter of the size, with lower cost and latency.

Inkling-Small is a July 30, 2026 open-weights multimodal autoregressive transformer model release from Thinking Machines Lab. It is a sparse Mixture-of-Experts model with 276B total parameters and 12B active parameters, accepts text, image, and audio inputs, generates text outputs, and supports up to a 1M-token context window.

Pricing

Official site · Jul 31, 2026, 7:32 AM
Input / 1M
$0.30
Output / 1M
$1.20
Cached input / 1M
$0.06

Serverless inference beta pricing per 1M tokens for Inkling-Small on Tinker; page notes beta is available for Inkling and Inkling-Small only and production use requires joining a waitlist.

Thinking Machines' official Tinker Models & Pricing page lists Serverless Inference (Beta) pricing for Inkling-Small and states all prices are per million tokens. The Inkling-Small row shows Prefill (Input) $0.30, cached input $0.06, and Sample (Output) $1.20 for the Tinker ID thinkingmachines/Inkling-Small:peft:262144:sampling-nvfp4.

View source

Model Intelligence

Context window
1,000,000 tokens
Arena ranking
41
Benchmarkable
Yes
Model level
release
Intelligence Index
41.2
Coding Index
52.9
GPQA
0.9
HLE
0.33
SciCode
0.49
LCR
0.69

Recent stories

1 linked story
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.