Skip to content
AI Primer
TOPIC50 stories

Model serving

Serving stacks and runtime systems for model inference.

RELEASE24th September
Quail open-sources MIT-licensed AI-SQL engine for LLM queries

Quail open-sourced an MIT-licensed engine that plans AI queries, batches inference, and reuses KV cache across filters and joins. Its authors report 1.84× faster execution than hand-tuned vLLM on 29 queries.

RELEASE21st September
Xiaomi releases open-weight MiMo-V2.6 models with 1M-token context

Xiaomi released Pro and Flash MiMo-V2.6 mixture-of-experts models with open weights and a 1M-token context window. The release includes an RL dashboard and day-one vLLM support, while training artifacts are planned.

RELEASE2w ago
DeepSeek releases 552B-parameter V4.1 Flash multimodal model

DeepSeek released V4.1 Flash, a 552B-parameter multimodal MoE model with 8B input and 16B output active parameters. It supports up to 1 million tokens of context, while an independent BridgeBench run used 23.5 million tokens on one task.

RELEASE2w ago
Cohere open-sources fused LLM decode kernel with 1.58x vLLM claim

Cohere released an open-source serving system that fuses the LLM decode step into one GPU kernel launch. On North Mini Code with one H100, it reports up to 1.58x vLLM performance at the tested batch size.

RELEASE3w ago
Perplexity open-sources Lily for Qwen3.6 inference on Apple silicon

Perplexity open-sourced Lily, a Rust and Metal engine for Qwen3.6-35B-A3B in Perplexity Computer's hybrid workflow. Perplexity reports 1.23× faster prefill and 1.35× faster decode on an M5 Max MacBook Pro.

RELEASE4w ago
Z.ai releases 743B-parameter GLM-5.3 open weights

Z.ai released GLM-5.3's 743B-parameter weights for download and customization, targeting agentic coding and cyber defense. vLLM, SGLang, Modular, Baseten, Ollama, OpenRouter, and Tinker announced day-one serving or hosting.

RELEASE4w ago
Tencent releases 770B-parameter Hy4 Preview open weights

Tencent released Hy4 Preview, a 770B-parameter mixture-of-experts model with 49B active parameters and a 1M-token context window. vLLM added day-zero support, while Cline, OpenCode Go, and Vercel AI Gateway made the model available.

WORKFLOW4w ago
Developer reports Qwen3.8-Flash-Next runs in 37GB on M4 Max

A developer reports streaming 60% of Qwen3.8-Flash-Next's experts from disk, running the full Q4 model in 37GB of RAM at 40 tokens per second. A separate H100 deployment reports 160 tokens per second with EAGLE-3.

NEWS4w ago
OpenAI reports Jalapeño delivers 1.5–1.9× more work per watt

OpenAI says Jalapeño delivered 1.5–1.9× more work per watt and 1.7–3.6× lower end-to-end latency than NVIDIA systems in its tests. The company plans to deploy the inference chip in its compute infrastructure by year-end.

NEWS4w ago
NVIDIA puts Groq 3 LPX into Vera Rubin production

Groq 3 LPX adds dedicated token generation to NVIDIA Vera Rubin systems, with Groq and Nebius among planned deployers. Artificial Analysis measured about 3,400 output tokens per second on Gemma 4 31B.

NEWS1mo ago
Tests link Ox Alpha to Zhipu GLM API routes and error codes

Researchers say malformed Ox Alpha requests exposed Zhipu-specific routes, error codes, and an internal class name. Independent vision comparisons also argue against speculation that the stealth model is Gemini.

NEWS1mo ago
Qwen 3.8 27B reaches 3,200 TPM at 262K context on two RTX 3090s

Community tests report Qwen 3.8 27B handling coding, OCR, and long-context workloads locally. One vLLM setup reached 3,200 tokens per minute at 262K context on two RTX 3090s without NVLink.

NEWS1mo ago
OpenCode Go revises limits after DeepSeek price increase

OpenCode says it revised Go limits after DeepSeek raised prices. Its operator is testing hosting configurations intended to bring DeepSeek service closer to its prior price point.

WORKFLOW1mo ago
Speculative decoding tests report acceptance drop from 0.71 to 0.18 after ~32K context

A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.

WORKFLOW1mo ago
Qwen 122B runs on older laptop with llama.cpp in user test

A LocalLLaMA post ran a 122B Qwen model on an older laptop with llama.cpp. The run had very long load and generation times, while another report put Qwen 3.6 35B at 21 tok/s on a Radeon 7600 after ROCm tuning.

NEWS1mo ago
Microsoft Copilot traces report 87% of LLM calls came from agents

A Microsoft Copilot trace analysis said 87% of LLM calls came from the agent, not direct user turns. Related posts warned token use and web requests can scale far faster than human prompt counts.

NEWS1mo ago
OpenCode users reportedly average $1.14/day on DeepSeek V4 Flash

OpenCode’s thdxr said Go users spent $1.14 per day on DeepSeek V4 Flash last week. Wafer added a fast OpenRouter route, while Nous extended a 90% discount for the 0731 model.

NEWS1mo ago
Together AI ranks first or tied first on 3 of 4 Kimi K3 provider benchmarks

Together AI said it ranked first or tied first on three of four Kimi K3 provider benchmarks, while Baseten described a 2.8T-parameter Blackwell GB300 serving stack. Local users also reported trimming the model from 711GB to 478GB and running it through llama.cpp RPC across clusters.

RELEASE1mo ago
Qwen3.8-Max launches on OpenRouter with 1M-token context

Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.

RELEASE1mo ago
DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context

Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

RELEASE1mo ago
MiniMax H3 releases Hugging Face weights and fal video endpoints

MiniMax H3 now has Hugging Face weights, fal endpoints, AI Toolkit LoRA support, and reported single-RTX-5090 local runs. MiniMax also said deployment in the US, EU, UK, and South Korea is available through formal authorization.

RELEASE1mo ago
Alibaba says Qwen3.8-Max open weights ship next week

Alibaba said Qwen3.8-Max left preview as a 2.4T-parameter MoE with 95B active parameters and $2/$6 per million-token pricing. Arena placed it on the Frontend Code Arena cost-performance frontier.

RELEASE1mo ago
MiniMax releases H3 open weights with day-zero vLLM-Omni support

MiniMax released H3 weights on Hugging Face for text-to-video, image-to-video, reference-to-video, and editing workflows. vLLM-Omni, ComfyUI, SGLang Diffusion, and fal added support at launch.

NEWS1mo ago
DeepSeek V4 Flash benchmarks show cheaper tokens but 3x SWE-Bench task cost

New tests showed DeepSeek V4 Flash as cheaper per token and faster on some serving paths. Ramp said it cost 3x more than GPT-5.6 Luna per SWE-Bench task because it used more turns.

RELEASE1mo ago
DeepSeek releases V4 Flash 0731 as MIT-licensed open weights

DeepSeek released V4 Flash 0731 with weights, a technical report, API access, 1M context, MoE routing, and low token prices. Its cited benchmarks show gains on Artificial Analysis, Terminal-Bench, Frontend Code Arena, and agent tests.

RELEASE1mo ago
Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim

Wafer listed Kimi K3 Fast on OpenRouter and Vercel AI Gateway. It claimed 172 output tokens/sec, 15.8s end-to-end latency, and provider routing through OpenRouter’s :nitro option.

RELEASE2mo ago
Together releases ThunderAgent for KV-cache scheduling in agent workflows

Together released ThunderAgent to schedule whole agent workflows instead of isolated requests during tool calls. Together reports up to 2.5x throughput and about 10x lower P50 latency under high concurrency.

RELEASE2mo ago
Kimi K3 launches across vLLM, SGLang, Ollama, and OpenRouter

Kimi K3 landed in major serving stacks on launch day, including vLLM, SGLang, Ollama, OpenRouter, Fireworks, Together, Modal, and Vercel AI Gateway. Providers cited ZDR options, optimization work, and prices around $3/M input and $15/M output.

RELEASE2mo ago
Google ships Gemini 3.6 Flash and 3.5 Flash-Lite to serving platforms

Gemini 3.6 Flash and 3.5 Flash-Lite went live on OpenRouter, Venice, Hyperbrowser, and Google surfaces. Early benchmarks show lower token use and cost, but mixed document and coding results.

RELEASE2mo ago
Martian launches Ship beta with 50% lower-cost inference target

Martian launched Ship, a beta endpoint that sits between an app and its reference model and routes each request through cheaper paths while aiming to preserve behavior. Martian says the beta targets 50% lower cost.

NEWS2mo ago
Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch

Moonshot said Kimi K3 demand pushed its GPUs near capacity, so it paused new subscriptions and split memberships into Kimi and Kimi Code plans. Users also reported slow serving and sold-out paid plans.

RELEASE2mo ago
Moonshot launches Kimi K3 with 2.8T parameters and 1M context

Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.

WORKFLOW2mo ago
LocalLLaMA users report near-6x Qwen 3.6 27B speedups with MTP

A LocalLLaMA benchmark on Qwen 3.6 27B and RTX 6000 PRO reports near-6x speedups from MTP, DFlash, and n-gram drafting. Related tests cover remote prefill, NUMA offload, GPU clock tuning, and llama.cpp Gemma 4 support.

NEWS2mo ago
Inkling adds early llama.cpp serving via 1-bit GGUF

Inkling's 1-bit GGUF ran in llama.cpp at 30–40 TPS, and TokenSpeed added day-zero support with a flat KV cache pool. Arena posts put Inkling #10 among open models in frontend code and text, while docs drew scrutiny.

RELEASE2mo ago
Thinking Machines releases Inkling: 975B open-weight multimodal MoE

Thinking Machines released Inkling with Apache 2.0 weights, 975B parameters, 41B active parameters, text/image/audio support, and up to 1M context. vLLM, SGLang, Modal, Databricks, and Vercel added day-zero support.

RELEASE2mo ago
vLLM v0.25.0 makes Model Runner V2 the default path for dense models

vLLM v0.25.0 made Model Runner V2 the standard dense-model execution path and removed legacy PagedAttention. The release also added parser, speculative decoding, distributed-serving, and security upgrades.

RELEASE2mo ago
Unsloth releases Qwen3.6 NVFP4 quants with claimed 2.5x GPU speedups

Unsloth released Qwen3.6 NVFP4 quants and claimed 2.5x GPU speedups, including 27B on 24GB VRAM. Follow-up notes warned vLLM users that Marlin or default backends can make W4A4 Qwen inference 2–2.5x slower.

RELEASE2mo ago
Tencent releases Hy3, a 295B MoE model under Apache license

Tencent released Hy3 with 21B active parameters, a 256K context window, BF16/FP8 weights, and day-one vLLM/SGLang support. Kilo Code, Nous Portal, and OpenRouter also made it free for limited windows.

RELEASE2mo ago
LongCat-2.0 opens MIT weights for 1.6T MoE with 1M context

Meituan released LongCat-2.0 weights and inference code under MIT, with Hugging Face, GitHub, ModelScope, GPU, and NPU paths. Analysts noted the ~48B-active MoE keeps attention shape while reducing zero-communication experts from 256 to 128.

WORKFLOW2mo ago
OpenRouter claims 24x inference-cost savings with MCP model routing

OpenRouter published an MCP workflow that it says cut inference costs 24x at comparable quality. The MCP lets the model choose providers using codebase context plus OpenRouter benchmark, aggregate-usage, and live-performance data.

NEWS2mo ago
Wafer reports GLM-5.2 hits 2,626 tok/s on MI355X

Wafer reported GLM-5.2 serving at 2,626 tok/s per MI355X node, and Together put it at 80% of Sonnet 5 capability for 20% of the price. Critics questioned whether public benchmark gains were overfit.

NEWS2mo ago
GLM-5.2 benchmarks at 97.6% tool-calling and 2,626 tok/s on MI355X

Kilo, Composio, Together, and Wafer posted GLM-5.2 measurements including 40/41 tool tasks, 7/10 code review, and 2,626 tok/s on MI355X. Try it for lower-cost coding and tool use, but validate cross-file reasoning and latency on your workload.

NEWS2mo ago
The Information reports OpenAI cuts inference costs by more than 50% on some models

Multiple summaries of The Information report said OpenAI found inference optimizations that more than halved costs on some existing models. If that holds, it changes the margin, pricing, and usage-limit math behind ChatGPT and API serving even before new model releases arrive.

RELEASE3mo ago
DeepSeek releases DSpark checkpoints for Qwen3 and Gemma-4

DeepSeek extended DSpark beyond V4 by publishing draft-model checkpoints for Qwen3 and Gemma-4 families and clarifying that DSpark targets higher-throughput serving by controlling verification cost. The release matters because speculative decoding is moving from papers into reusable open checkpoints.

RELEASE3mo ago
DeepSeek V4-Pro benchmarks at ~90 tok/s after DSpark rollout

Independent measurements after DSpark put DeepSeek V4-Pro around 90 tok/s and cut one run from 214s to 116s. The gain matters because it lowers serving cost, though tuning details and memory overhead are still unclear.

RELEASE3mo ago
DeepSeek releases DeepSpec and DSpark for speculative decoding on V4 checkpoints

DeepSeek open-sourced DeepSpec, a codebase for training and evaluating draft models for speculative decoding, alongside the DSpark decoding module for V4 checkpoints. It matters because inference teams get a new open stack for improving draft-model quality and decode throughput beyond earlier MTP-style baselines.

RELEASE3mo ago
Vercel AI Gateway adds GLM-5.2 Fast at 150-250 tok/s

Vercel and Wafer launched a serverless GLM-5.2 endpoint on AI Gateway with 1M context and published pricing. Teams get a high-throughput open-model option inside an existing gateway instead of managing GLM inference directly.

NEWS3mo ago
GLM-5.2 adds Perplexity Agent API and Droid support on Baseten at >280 TPS

GLM-5.2 added Perplexity Agent API, Droid, and more hosting options, while Baseten reported over 280 TPS and sub-0.8s TTFT. Builders should watch the cost and benchmark data as it moves into production agent stacks.

RELEASE3mo ago
Morph supports Qwen, GLM-5.2, MiniMax M3, DeepSeek v4 with 20-35% higher code acceptance

Morph said its code-serving stack now exposes Qwen, GLM-5.2, MiniMax M3, and DeepSeek v4 with code-tuned speculative decoding. It claims 20-35% higher acceptance than Eagle 3.1 or DFlash, plus kernels for cheaper hardware.

NEWS3mo ago
GLM-5.2 ships to BrowserCode, Hyper, OpenCode, and Together in 3 days

BrowserCode, Hyper, OpenCode, Together, and other vendors added GLM-5.2 soon after release. That turns the open model into a deployable option across coding, browser automation, and hosted chat.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.