Skip to content
AI Primer
release

Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim

Wafer listed Kimi K3 Fast on OpenRouter and Vercel AI Gateway. It claimed 172 output tokens/sec, 15.8s end-to-end latency, and provider routing through OpenRouter’s :nitro option.

6 min read
Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim
Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim

TL;DR

  • Wafer put Kimi K3 Fast on both OpenRouter and Vercel AI Gateway, with the OpenRouter launch in wafer_ai's post and the Vercel launch in wafer_ai's Vercel post.
  • OpenRouter's fast path is :nitro: OpenRouter's post said the suffix routes across providers for highest token throughput.
  • Wafer claimed 172 output tokens/sec and 15.8s end-to-end response time for a dedicated endpoint, with the numbers coming from wafer_ai's Artificial Analysis screenshot.
  • The cross-platform availability claim covered Artificial Analysis, OpenRouter, and Vercel AI Gateway in wafer_ai's follow-up.
  • The operational backdrop is huge: ThursdAI's episode thread put Kimi K3 at 2.8T parameters and 1.56 TB to serve, with tokenization becoming visible at 100K to 1M-token prompts.

OpenRouter's Kimi K3 page now exposes Balanced, Nitro, and Exacto routing modes, plus provider tables for P50 latency, throughput, uptime, cache-hit pricing, and traffic share. Vercel's Kimi K3 Fast page gives the separate model slug moonshotai/kimi-k3-fast, a 1,000,000-token context window, a 1,000,000-token max output, and fast-tier pricing at $4.50 input and $22.50 output per 1M tokens. Moonshot's Kimi K3 model card puts the underlying model at 2.8T parameters, Kimi Delta Attention, Attention Residuals, native vision, and 1M context.

OpenRouter Nitro

OpenRouter expressed the fast lane as routing, not a new base-model name. The important string is :nitro, which OpenRouter said automatically routes across providers for highest token throughput.

The launch sequence was tight:

  • Wafer announced Kimi K3 Fast on OpenRouter in its launch post.
  • Wafer repeated that Kimi K3 Fast was live on OpenRouter in a follow-up.
  • OpenRouter's current model page describes three routing modes: Balanced for price plus speed, Nitro for fastest, and Exacto for highest tool-calling accuracy.
  • The same page lists OpenAI-compatible chat completions and responses endpoints for moonshotai/kimi-k3.

OpenRouter's page also reports prompt-cache economics. Its rolling 30-day effective pricing table showed weighted average input cost at $0.697 per 1M tokens after cache hits, while list input prices for common providers sat around $3.00 per 1M.

Vercel Fast

Vercel exposed the fast path as a separate Gateway model slug. The public model page lists moonshotai/kimi-k3-fast, type chat, 1M context, 1M max output, and $4.50 input plus $22.50 output per 1M tokens.

The provider metadata is messy in the way launch-day gateway metadata often is. Vercel's model page read during research listed Fireworks and Morph as providers, while wafer_ai's post said Kimi K3 Fast by Wafer was live on Vercel AI Gateway and told users to pick Wafer as provider.

That split makes the API shape worth naming: OpenRouter uses a Nitro routing suffix, while Vercel uses a Fast model slug.

172 output tokens per second

Wafer's strongest number came from its Artificial Analysis screenshot: 172 output tokens/sec and 15.8s end-to-end response time.

The screenshot's highlighted provider comparison:

  • Speed: Wafer Fast 172 output tok/s, Fireworks 164, Makora 137, Databricks 123, Modal 104.
  • End-to-end response time: Wafer Fast 15.8s, Fireworks 16.4s, Makora 19.6s, Databricks 21.5s, Modal 25.8s.
  • Blended price: Kimi, Fireworks, Modal, Together AI, DigitalOcean, Parasail, Makora, and Databricks at $2.3 per 1M tokens, Wafer Fast at $3.5, Nebius at $4.2.

The public Artificial Analysis provider page read during research showed a different Kimi K3 provider set, with Makora at 165.0 t/s, Fireworks at 163.3 t/s, and Databricks at 136.9 t/s. The same page says provider performance varies over time because of infrastructure changes, load balancing, and updates.

Fast-tier pricing

The fast tier carries a clear premium in the launch screenshots. OpenRouter's table had fast providers at $4.50 input, $22.50 output, and $0.45 cache-read per 1M tokens, while standard providers sat at $3.00 input, $15.00 output, and $0.30 cache-read.

The screenshot's point-in-time OpenRouter table showed:

  • Fireworks Fast: $4.50 input, $22.50 output, $0.45 cache read, 1.07s latency, 87 tps.
  • Modal: $3.00 input, $15.00 output, $0.30 cache read, 1.19s latency, 82 tps.
  • Wafer Fast: $4.50 input, $22.50 output, $0.45 cache read, 1.60s latency, 79 tps.
  • Wafer: $3.00 input, $15.00 output, $0.30 cache read, 3.45s latency, 59 tps.
  • Fireworks: $3.00 input, $15.00 output, $0.30 cache read, 1.81s latency, 48 tps.

AAAzzam's OpenRouter listing post framed the race from the other side: it was his first time listing a model on OpenRouter, with a public nudge about releasing a Fast version of Kimi K3.

Kimi K3 serving shape

The reason a provider race formed around Kimi K3 is simple: the open weights are large enough that the model is technically downloadable and operationally nontrivial.

Moonshot's Hugging Face model card describes Kimi K3 as:

  • 2.8T total parameters.
  • 104B active parameters per token, according to independent release analysis.
  • Kimi Delta Attention plus Attention Residuals.
  • Native vision.
  • 1M-token context.
  • Stable LatentMoE with 16 of 896 experts activated per token.

A Backgrind release analysis measured the released checkpoint at 1.56 TB on disk and described the API versus self-hosting split bluntly: aggregators make sense for fallback and one-bill access, while self-hosting mostly buys auditability, private deployment, and modification rights.

Morph and SGLang

Morph had already tied Kimi K3 Fast to inference-stack work rather than only gateway placement. Its announcement said Morph was working with SGLang and LMSYS on open-model disaggregated inference, with Kimi K3 Fast served at up to 100 tokens/sec through OpenAI and Anthropic-compatible APIs.

That matters because Vercel's public Kimi K3 Fast page listed Morph as one of its providers during research. Morph's post also described the Kimi K3 Fast work as the beginning of a deeper collaboration around low-level optimization and high-performance open-model serving.

WeirdML and tokenizer drag

Kimi K3's speed story is not only decode throughput. ThursdAI's episode thread called tokenization the funny bottleneck: with 100K to 1M-token inputs and a hot prefix cache, tokenization can materially affect time to first token.

TeortaxesTex put Kimi K3 near Opus 4.8 (xhigh) level on WeirdML and described it as a large jump from GLM 5.2. The same thread noted greedy decoding as a strong baseline, a reminder that provider benchmarks, model evals, and decoding settings are measuring different parts of the stack.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR2 posts
OpenRouter Nitro2 posts
172 output tokens per second1 post
Fast-tier pricing1 post
Kimi K3 serving shape1 post
Morph and SGLang1 post
Share on X