MODEL3 stories
Inkling
Stories, products, and related signals connected to this tag in Explore.
Stories
Filter storiesRELEASE4w ago
Thinking Machines releases Inkling-Small 276B open-weight MoE
Thinking Machines released Inkling-Small, a 276B-parameter MoE with 12B active parameters, multimodal inputs, and a 1M-token context window. Providers added day-zero vLLM, SGLang, Modal, and gateway support.
NEWS1mo ago
Inkling adds early llama.cpp serving via 1-bit GGUF
Inkling's 1-bit GGUF ran in llama.cpp at 30–40 TPS, and TokenSpeed added day-zero support with a flat KV cache pool. Arena posts put Inkling #10 among open models in frontend code and text, while docs drew scrutiny.
RELEASE1mo ago
Thinking Machines releases Inkling: 975B open-weight multimodal MoE
Thinking Machines released Inkling with Apache 2.0 weights, 975B parameters, 41B active parameters, text/image/audio support, and up to 1M context. vLLM, SGLang, Modal, Databricks, and Vercel added day-zero support.