Skip to content
AI Primer
update

Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch

Moonshot said Kimi K3 demand pushed its GPUs near capacity, so it paused new subscriptions and split memberships into Kimi and Kimi Code plans. Users also reported slow serving and sold-out paid plans.

6 min read
Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch
Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch

TL;DR

The Kimi API quickstart makes the scale problem concrete: 2.8T parameters, 16 of 896 experts active, 1M context, default max thinking, OpenAI-compatible calls, and full weights promised by July 27. Artificial Analysis scored it 57, behind Fable 5 and GPT-5.6 Sol, while AlphaSignal put the same launch window into held-out repair tasks and got 702 seconds per attempt. The strangest product detail was musical: Moderato, Allegretto, Allegro, and Vivace were all sold out in one Kimi pricing screenshot.

The subscription freeze

Moonshot said Kimi K3 had received more demand than expected and that GPUs were close to the limits of current capacity. The company paused new subscriptions, prioritized compute for existing subscribers, and said new spots would reopen in batches.

The same notice said Kimi would split membership into two plans: Kimi Membership for Kimi Web, App, and Work, plus Kimi Code Membership for coding workflows. TestingCatalog framed the move as a compute waitlist for K3.

Sold-out plans

In bridgemindai's screenshot, annual billing showed four sold-out tiers: Moderato at $15/month, Allegretto at $31/month, Allegro at $79/month, and Vivace at $159/month. The listed features bundled agent credits, Docs/Sheets/Slides integrations, Deep Research, Websites Deploy, Agent multi-tasking, Kimi Code credits, and 1M-token conversations at the top tier.

In Scott Stevenson's screenshot, Kimi answered a prompt with, “Too many people are chatting with Kimi right now,” then pushed users toward a dedicated priority queue. Stevenson later posted a Vivace purchase screen showing the C$249 monthly and C$2,799.99 annual options greyed out as sold out.

Serving bottleneck

On OpenRouter, bridgemindai's provider snapshot listed Moonshot AI as the only provider, with $3/M input, $15/M output, $0.30/M cache read, 11.22s latency, 16 tok/s throughput, and 99.97% uptime.

bridgemindai's later chart put average throughput at 24 tok/s and average latency at 5.78s. In deedydas's compute thread, K3 was already #10 on OpenRouter at roughly 140B tokens/day, with throughput down from 30 to 13 tok/s, end-to-end latency up to 72s, and time to first token above 20s.

Task runners saw the same failure mode. mattlam_ said Kimi K3 runs were timing out on tougher OpenBench tasks with a 20-minute timeout, and teortaxesTex posted a “Task paused due to system peak” message.

Benchmark split

K3's launch numbers were strong enough to make the capacity crunch painful.

  • Frontend Code Arena: Kolt Regaskes' follow-up said Kimi K3 reached #1 with 1679 points, ahead of Claude Fable 5.
  • DeepSWE: Kolt Regaskes' chart put K3 max at 69% ±5, $4.65 per task, 81k output tokens, and 98 steps.
  • Artificial Analysis: Artificial Analysis scored K3 at 57, comparable to Opus 4.8 and GPT-5.5, behind Fable 5 and GPT-5.6 Sol.

AlphaSignal measured a rougher agentic repair workload. AlphaSignal's report used 13 planted-bug tasks, seven models, network-off Docker sandboxes, and held-out tests only at scoring; K3 resolved 53 of 67 attempts, finished last of seven, cost $0.186 per successful fix, and averaged 702 seconds per attempt.

Hands-on reports matched the wall-clock caveat. kunchenguid said K3 understood intent and could diagnose problems, but felt very slow, burned one-third of a $40 plan's five-hour limit after a few prompts, and missed firstmate system-prompt instructions that other frontier models followed.

The token bill

The Kimi K3 pricing docs list three API prices per 1M tokens: $0.30 cache-hit input, $3.00 cache-miss input, and $15.00 output, with a 1,048,576-token context window. The Kimi API quickstart says context caching is automatic when the long prefix remains unchanged.

On DeepSWE economics, Rohan Paul's chart said $100 bought 14.7 solved K3 rollouts versus 5.3 for Fable 5, while K3's median rollout took 66 minutes versus Fable's 21 minutes.

theo argued K3 is “an incredible model” but not an incredible value in most tasks: K3 is about half the per-token price of GPT-5.6 Sol, while Sol uses roughly half as many tokens and runs at about twice the tokens per second. Emad Mostaque pointed to the other side of the bill, saying K3 is optimized for high cache-hit rates and could run cheaper on large-memory B300-class clusters that Moonshot does not have.

Open weights and partner serving

The Kimi API quickstart says Moonshot is working with inference partners and open-source maintainers, with full weights due by July 27. bigeagle_xd said the team was working “extremely hard” with partners to open the weights sooner.

Providers immediately treated the weights as an infrastructure event. TogetherCompute said users would be able to deploy K3 in a preferred region or use a domestic serverless endpoint once it was available on its platform, and Yuchenj_UW said Databricks was excited to serve K3 soon.

The hardware estimates were not hobbyist-friendly. deedydas estimated a minimum $500k for the 8 B300s needed to serve even quantized K3, with a roughly $4M GB300 NVL72 rack as the more recommended setup, while arankomatsuzaki's memory-fit chart marked H200 insufficient even at 4-bit.

Kimi Code access paths

The membership split had already shown up in developer surfaces. superdoteng's screenshot showed a Kimi Code TUI provider running K3 with “thinking: max” and a 0/1M context indicator, available to Moderato users or above.

ai_for_success showed the Kimi Code API key flow into Hermes: create a key in the Kimi Code console, paste it into Hermes settings under Providers, and select K3. Cline offered Kimi K3 through ClinePass with a $1.99 promo, and Cline's follow-up said the same OpenAI-compatible API could be used outside Cline.

Moonshot also opened team purchasing before the pause became the story. Moonshot's Kimi Business post listed Kimi Business at a five-seat annual minimum with enterprise-grade data privacy, bank transfer, self-service invoicing, dedicated technical support, and Kimi Code credits.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 7 threads
The subscription freeze1 post
Sold-out plans1 post
Serving bottleneck3 posts
Benchmark split1 post
The token bill1 post
Open weights and partner serving3 posts
Kimi Code access paths3 posts
Share on X