Inkling
Open weights, ready to tinker
Inkling is an open-weights, general-purpose multimodal autoregressive transformer model from Thinking Machines Lab. The official model card states it accepts text, image, and audio inputs and generates text outputs; it is a sparse Mixture-of-Experts model with 975B total parameters, 41B active parameters, Apache 2.0 licensing, and support for up to a 1M-token context window.
Pricing
Model Intelligence
Recent stories
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.
Inkling's 1-bit GGUF ran in llama.cpp at 30–40 TPS, and TokenSpeed added day-zero support with a flat KV cache pool. Arena posts put Inkling #10 among open models in frontend code and text, while docs drew scrutiny.
Thinking Machines released Inkling with Apache 2.0 weights, 975B parameters, 41B active parameters, text/image/audio support, and up to 1M context. vLLM, SGLang, Modal, Databricks, and Vercel added day-zero support.