Kimi K3 ranks No. 1 on Arena Frontend Code leaderboard
Posts put Kimi K3 first on Arena's Frontend Code leaderboard but 4.37-5.29 months behind U.S. frontier models in one public estimate. Other evidence cited strong DeepSWE cost-performance and cybersecurity results.

TL;DR
- Kimi K3 took the public frontend crown: arena's announcement put it at 1679 points, first on Frontend Code Arena and ahead of Claude Fable 5.
- The model package is huge and still waiting on downloadable weights: the spec post says K3 has 2.8T parameters, a 1M context window, and open weights due by July 27.
- Coding evals split hard by task shape: one benchmark roundup put K3 at 69% on DeepSWE for $4.65 per task, while AlphaSignalAI's repair run had it last of 7 models on a held-out bug-fix harness.
- The cost story is workload-dependent: Cline's bug test found Kimi 2.3x cheaper but Fable 3.4x faster, and theo's cost note argued GPT-5.6 Sol often erases K3's per-token discount by using about half as many tokens.
- Cyber is the policy pressure point: the internal cyber eval post called K3 top-tier for cybersecurity while saying Sol was stronger and Fable refused the run.
Moonshot's Kimi K3 docs describe a 2.8T flagship built on Kimi Delta Attention and Attention Residuals, with native vision and a 1M-token context window. The K3 pricing page exposes $0.30 cached input, $3 fresh input, and $15 output per million tokens, plus always-on reasoning with reasoning_effort currently set to max in the API. The dynamic tools docs are agent bait: K3 can lazy-load tool definitions instead of stuffing every MCP tool into the initial prompt.
K3 package
Moonshot's docs call K3 its most capable model to date and say the full weights will be released by July 27. The current public surface is hosted Kimi, API, Kimi Work, Kimi Code, and partner routes, while the open-weight claim is still pending the actual files.
The developer-relevant package:
- 2.8T parameters, described by Moonshot as a 3T-class open model.
- 1,048,576-token context window.
- Native visual understanding.
- Kimi Delta Attention plus Attention Residuals.
- Automatic context caching, tool calls, JSON mode, structured output, partial mode, and internet search support.
- API pricing at $0.30 cached input, $3 fresh input, $15 output per 1M tokens.
K3 also adds deferred tools. bigeagle_xd's tool note said dynamically loaded tools reduce initial context length for agents with many tools or MCPs, and a follow-up said newly loaded tools are appended to the context rather than prepended.
Frontend Code Arena
Arena said Kimi-K3 is the first Chinese model to take the lead over U.S. models in Frontend Code Arena. arena's head-to-head post then published identical Kimi K3 and Fable 5 prompt pairs across ramen sites, dashboards, booking systems, physics tables, orbit consoles, and other UI tasks.
The public result was clean:
- Kimi-K3: 1679 points, No. 1 on Frontend Code Arena, per arena's announcement.
- Claude Fable 5: surpassed by K3 in Arena's post.
- The same Arena thread linked 16 Kimi and Fable outputs for side-by-side inspection, via the prompt-pair list.
Frontend is the shareable benchmark class: screenshots, taste, layout, animation, and working UI primitives travel faster than backend diffs. That is why this leaderboard did more to move the conversation than most aggregate scores.
Coding agent split
DeepSWE and real repair harnesses told different stories.
- DeepSWE: K3 scored 69% at $4.65 per task, versus Sol at 73% and $8.39, and Fable 5 at 70% and $21.63, according to one benchmark roundup.
- K2.7 to K3 jump: the DeepSWE explainer said Kimi moved from 31 points with K2.7-code to 69 points with K3, a 38-point gain.
- Artificial Analysis Coding Agent Index: the coding index post gave K3 a score of 57, joint No. 5, with 84% on Terminal-Bench v2, 64% on DeepSWE, and 23% on SWE-Atlas-QnA.
- Cline real bug: Cline's comparison said both K3 and Fable fixed the issue, Kimi used 1.2M tokens versus Fable's 730K, Fable finished in 3.5 minutes versus Kimi's 12 minutes, and Kimi cost $0.92 versus $2.13.
- AlphaSignalAI repair harness: the method note described 13 planted-bug tasks, network-off Docker sandboxes, and held-out scoring; K3 finished 53 of 67 attempts, 79%, while Sol hit 70 of 70.
AlphaSignalAI's useful sentence was the boring one: frontend Elo and held-out repair are different jobs. one AlphaSignalAI reply framed 1679 as frontend preference and 79% as whether a fix survives tests.
Cost and latency
Moonshot paused new subscriptions after demand pushed close to capacity, while keeping existing subscribers active. one latency follow-up claimed K3 later ran 100% faster after Moonshot stopped selling new plans.
K3's bargain story has three separate variables:
- Sticker price: $3 input and $15 output per 1M fresh tokens, with cached input at $0.30, per Moonshot's pricing docs.
- Reasoning verbosity: theo's note said GPT-5.6 Sol costs about twice as much per token but often uses half as many tokens.
- Serving speed: BridgeMind's latency post measured 24 tokens per second and nearly 6 seconds before first token; an earlier BridgeMind post reported 16 tokens per second and 11 seconds of latency.
OpenRouter demand was already large enough to stress the serving layer. Deedy Das's compute thread claimed K3 reached roughly 140B tokens per day on OpenRouter within two days, with throughput falling from 30 tokens per second to 13 and end-to-end latency rising to 72 seconds.
A small Reddit thread captured the skeptical version of the same point: the r/OpenaiCodex post complained that K3's output-token appetite made it more expensive than U.S. models in practice, while granting that the UI design was strong.
Cyber gap
Cyber is where K3 stopped being a coding leaderboard story and became a governance story.
The internal cyber eval post said K3 was top-tier, Sol was a leap ahead at higher cost, and Fable refused the run. cramforce's private benchmark gave the more operational breakdown: Sol had the best recall and precision at more than 7x the cost of the runner-up, K3 had the best price/recall tier, GLM-5.2 was 40% cheaper than K3 at good recall, and Fable had a 100% refusal rate.
The unresolved public benchmark is AISI. one CyberGym note pointed out that K3's launch benchmarks did not include a CyberGym score, and the UK AI Security Institute's open-weight cyber report found GLM-5.2 and DeepSeek V4-Pro only 4 to 7 months behind closed models on cyber, down from 6 to 10 months through much of 2025.
Deployment surface
K3 spread through coding tools before the weights dropped.
- Vercel AI Gateway added the model slug
moonshotai/kimi-k3, according to Vercel's gateway post. - OpenCode added K3 through Moonshot auth and
/models, per the OpenCode setup post. - Kimi Code got a native TUI provider after earlier OpenCode and Pi access, per the Kimi Code TUI post.
- ClinePass added K3 in a $9.99/month plan for discounted open-weight model access, according to Cline's ClinePass post.
- Together said it would serve K3 natively starting July 27, in Vipul Ved Prakash's provider post.
- Pi highlighted native multimodal support, long-horizon workflows, and faster million-token context in Pi's K3 post.
The routing hacks arrived too. AI Builder Club's Codex setup used CC Switch to translate Codex's Responses API calls into Kimi's Chat Completions format locally, and one Claude Code config pointed Anthropic model environment variables at kimi-k3[1m].