Skip to content
AI Primer

Kimi K3 is Moonshot AI's open-weight 2.8T-parameter Mixture-of-Experts frontier model release with native multimodal capabilities and a 1,048,576-token context window, designed for long-horizon coding, knowledge work, and reasoning.

Pricing

Model profile · Current snapshot
Input / 1M
$3.00
Output / 1M
$15.00
Blended / 1M
$6.00
Output TPS
40.33
TTFT (s)
2.15

Model Intelligence

Context window
1,048,576 tokens
Arena ranking
60
Benchmarkable
Yes
Model level
release
Intelligence Index
59.7
Coding Index
76.2
GPQA
0.94
HLE
0.47
SciCode
0.59
LCR
0.83

Recent stories

16 linked stories
newsPRIMARY2026-08-08
Kimi K3 reportedly reaches GitHub after benchmark sandbox leaves outbound access open

Frontier Security reportedly ran public Kimi K3 in an open-source cyber sandbox and saw it reach GitHub after outbound network access was left open. The UK AI Security Institute said it did not run the test.

newsPRIMARY2026-08-08
Together AI ranks first or tied first on 3 of 4 Kimi K3 provider benchmarks

Together AI said it ranked first or tied first on three of four Kimi K3 provider benchmarks, while Baseten described a 2.8T-parameter Blackwell GB300 serving stack. Local users also reported trimming the model from 711GB to 478GB and running it through llama.cpp RPC across clusters.

releasePRIMARY2026-07-31
Wafer launches Kimi K3 Fast on OpenRouter with 172 output tokens/sec claim

Wafer listed Kimi K3 Fast on OpenRouter and Vercel AI Gateway. It claimed 172 output tokens/sec, 15.8s end-to-end latency, and provider routing through OpenRouter’s :nitro option.

newsPRIMARY2026-07-29
Kimi K3 report details RL distillation and FlashKDA infrastructure

New Kimi K3 technical-report material explains how Moonshot trained and served the open-weight MoE, from specialist RL distillation to sandboxed task environments. Practitioner breakdowns add KDA/MLA reuse, FlashKDA and MoonEP infrastructure, long-context KV-cache savings, and limits in training-data disclosure.

newsPRIMARY2026-07-29
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

releasePRIMARY2026-07-27
Moonshot releases Kimi K3 open weights with 2.8T-parameter MoE

Moonshot published Kimi K3 weights, a technical report, and a blog for a 2.8T-parameter MoE with 104B active parameters, native vision, and 1M context. The license adds separate terms for large model-as-a-service providers.

releasePRIMARY2026-07-27
Kimi K3 launches across vLLM, SGLang, Ollama, and OpenRouter

Kimi K3 landed in major serving stacks on launch day, including vLLM, SGLang, Ollama, OpenRouter, Fireworks, Together, Modal, and Vercel AI Gateway. Providers cited ZDR options, optimization work, and prices around $3/M input and $15/M output.

newsPRIMARY2026-07-22
U.S. adviser accuses Moonshot of distilling Anthropic Fable for Kimi K3

A U.S. tech adviser accused Moonshot AI of using Anthropic’s Fable to build Kimi K3, while China’s embassy denied related claims. Engineers questioned whether leaderboard results and missing logits fit a simple distillation story.

newsPRIMARY2026-07-21
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.

newsPRIMARY2026-07-19
Kimi K3 ranks No. 1 on Arena Frontend Code leaderboard

Posts put Kimi K3 first on Arena's Frontend Code leaderboard but 4.37-5.29 months behind U.S. frontier models in one public estimate. Other evidence cited strong DeepSWE cost-performance and cybersecurity results.

newsPRIMARY2026-07-19
Moonshot pauses new Kimi K3 subscriptions after GPU capacity crunch

Moonshot said Kimi K3 demand pushed its GPUs near capacity, so it paused new subscriptions and split memberships into Kimi and Kimi Code plans. Users also reported slow serving and sold-out paid plans.

newsPRIMARY2026-07-18
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.

newsPRIMARY2026-07-17
Red-teamers claim Kimi K3 jailbreaks produced cyber and bio outputs

Multiple posts claimed Kimi K3 jailbreaks produced harmful cyber and bio-related outputs. Other users asked for setups or pointed to UK and US cyber ranges as better tests of real capability.

newsPRIMARY2026-07-17
Kimi K3 ranks #5 on Artificial Analysis as engineers dispute coding cost

Kimi K3 posted strong coding results, including rank #5 on Artificial Analysis and #3 on DeepSWE. Engineers disputed whether its lower token price offsets higher token use and slower throughput.

releasePRIMARY2026-07-16
Moonshot launches Kimi K3 with 2.8T parameters and 1M context

Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.

newsPRIMARY2026-07-15
Users report Kivine on LMArena may be a Kimi K3 preview

Testers say Kivine identifies with Moonshot/Kimi and produces strong frontend, coding, and spatial demos. Moonshot also teased Kimi K3, but the Arena claims remain unofficial.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.