Skip to content
AI Primer
update

Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

AlphaSignal's repair harness put Kimi K3 last at 53/67, while other tests ranked it high on DeepSWE, Vibe Code Bench, legal work, cyber, and CUDA kernels. Cost often beat Fable, but speed lagged.

6 min read
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness
Kimi K3 benchmarks last at 53/67 in AlphaSignal repair harness

TL;DR

  • Kimi K3 shipped as a 2.8T, 1M-context, native-vision model with open weights promised by July 27, according to Kimi_Moonshot's launch post.
  • The coding story split hard by workload: K3 topped Frontend Code Arena, while AlphaSignalAI's repair run put it last at 53/67 on held-out bug fixes.
  • The cost story is workload-dependent: ArtificialAnlys's index post measured $0.94 per Intelligence Index task, while Cline's real-bug run found Kimi cheaper than Fable but much slower.
  • The main operational gotcha is state handling, because AlphaSignalAI's follow-up says K3 needs full thinking history each turn and Kimi's own docs currently expose only max reasoning.

Moonshot's official tech blog says K3 built a MiniTriton compiler, designed a chip in a 48-hour autonomous run, and used 2,800+ web fetches for an ASIC industry report. The API docs hide the harness contract: multi-turn callers must append the complete assistant message, not just content. Simon Willison's pelican test cost 25 cents because K3 spent 13,241 reasoning tokens on one SVG.

What shipped

Benchmarks that moved

First-party

Third-party evaluators

Customer-reported

Where it regressed

AlphaSignalAI got the loudest counter-result: K3 finished last in a repair harness even while other boards ranked it near the top.

Ofir Press, one of the ProgramBench authors, said Moonshot used an average implementation metric rather than the benchmark's recommended fully-implemented-program metric. That can turn "90% of every program" into a high score even when zero programs fully pass, OfirPress's ProgramBench caveat argued.

Moonshot also acknowledged a product gap: scaling01's screenshot of the Kimi blog quoted Kimi's own line that K3 still has a noticeable user-experience gap versus Claude Fable 5 and GPT-5.6 Sol.

Under the hood

K3's model card is really a systems story.

Vibe Check

Hands-on reports mostly agreed on the shape: strong output, expensive thinking, uneven control.

  • K3 was "very, very slow" in a firstmate session, burned a third of a 5-hour plan in a few prompts, and missed system-prompt instructions that other frontier models followed, according to kunchenguid's firstmate report.
  • A 1.5 MB FrankenGraphDB plan review produced findings that Fable later judged substantially correct, though about a third of remedies needed correction, according to doodlestein's meta-review thread.
  • A simple TUI hover-color fix went to Sol for $0.30, while Kimi reached $1.00 and started reading the database before being interrupted, according to thdxr's task report.
  • Cline's traces suggested Kimi is RL-trained to spend more tokens thinking and verifying before completion, according to Cline's real-bug run.
  • OpenRouter users saw the operational side immediately: one rate-limit complaint could not get two successful requests in a row, and OpenRouter's K3 page warned that upstream capacity was limited and frequent 429s were possible.
  • Simon Willison's pelican writeup found the same always-on-max pattern in miniature: 16,658 output tokens for a toy SVG, including 13,241 reasoning tokens.

Where it shows up

K3 became a day-one integration race.

  • Kimi's own surfaces shipped first: Kimi.com, Kimi Work, Kimi Code, and API access, according to Kimi_Moonshot's launch post.
  • Vercel AI Gateway added moonshotai/kimi-k3, and Vercel's Kimi Code plugin note said Kimi Code could draw on Vercel platform knowledge for Next.js, AI SDK, and Vercel Functions.
  • Hermes Agent added Kimi through Nous Portal, Kimi Direct, and OpenRouter, according to Teknium's Hermes post.
  • Cline put K3 into ClinePass, with a $1.99 promo for users installing via npm i -g cline, according to Cline's subscription post.
  • OpenCode made K3 available to OpenCode Go users before negotiating a discount, according to OpenCode's rollout note.
  • Pi added K3 with native multimodal support, long-horizon self-evolving workflows, and faster million-token context, according to Pi's integration post.
  • AI/ML API put K3 and Claude Fable 5 behind the same key, SDK, billing, and OpenAI-compatible endpoint, according to TestingCatalog's AI/ML API post.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 7 threads
TL;DR4 posts
What shipped2 posts
Benchmarks that moved9 posts
Where it regressed6 posts
Under the hood3 posts
Vibe Check4 posts
Where it shows up7 posts
Share on X