Skip to content
AI Primer
update

Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

Artificial Analysis reports Kimi K3 averages 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task. Kilo also found UI-build outputs close to Claude Fable 5 at 29% of the cost.

6 min read
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase
Artificial Analysis reports Kimi K3 averages 56.4 minutes on AA-Briefcase

TL;DR

  • Kimi K3 is now second on AA-Briefcase: Artificial Analysis reported a 1543 Elo, behind Claude Fable 5 at 1574 and up +727 over Kimi K2.6.
  • The score came with a heavy run profile: Artificial Analysis' time breakdown measured 56.4 minutes, 83 turns, and 120k output tokens per AA-Briefcase task.
  • Frontend work was the cleanest cost story: Kilo's head-to-head found ten one-shot UI builds looked eerily similar to Fable 5 at 29% of the cost.
  • Launch-week serving became the limiting system: Kimi paused new subscriptions after GPU demand hit capacity, while bridgemindai measured one-provider OpenRouter serving at 16 tps and 11.22s latency.

Kimi's own tech blog says K3 launched at max thinking effort by default, with low and high effort modes coming later. Artificial Analysis turned that into the most useful bill of materials: tokens, turns, minutes, and dollars. Kilo found the opposite side of the trade: UI outputs converged on Fable 5's structure at much lower token pricing.

AA-Briefcase

The full Artificial Analysis article defines AA-Briefcase as private, realistic knowledge-work tasks across thousands of complex input files. Tasks require deliverables such as spreadsheets, presentations, and UI mock-ups, then roll correctness, analytical quality, and presentation quality into one Elo.

Artificial Analysis reported:

  • AA-Briefcase Elo: Kimi K3 1543, Claude Fable 5 1574, GPT-5.6 Sol 1501.
  • Generation jump: Kimi K3 1543, Kimi K2.6 816, a +727 gain.
  • Rubric pass rate: Kimi K3 51%, Claude Fable 5 56%.
  • Analytical quality Elo: Kimi K3 1754, Claude Fable 5 1744.
  • Presentation Elo: Kimi K3 1471, below GPT-5.6 Sol at 1660 and Claude Opus 4.8 at 1492.

Christmas came early for open-weight leaderboard nerds, but the benchmark split is awkward: K3 looks close to Fable on analysis and much weaker on polish.

Cost per task

The cost section is the part to bookmark. K3's AA-Briefcase score came from a long, turn-heavy run rather than a cheap flash of intelligence.

Artificial Analysis measured:

  • Average cost: $10.57 per task, below Claude Sonnet 5 at $14.43 and Claude Fable 5 at $22.30.
  • Relative cost: GPT-5.6 Sol trails K3 by 42 Elo while costing about 50% less at $5.32 per task.
  • Wall-clock time: 56.4 minutes per task, about 2.5x Claude Fable 5 and 3.8x Grok 4.5 high.
  • Output volume: 120k output tokens per task, up from 42k for Kimi K2.6.
  • Turns: 83 turns per task, up from 54 for Kimi K2.6.
  • Token pricing: $3 per 1M input tokens, $15 per 1M output tokens, with a 90% discount for cached tokens.

That is the K3 launch-week shape in one chart: near-frontier work, slow enough to make agent loops feel like batch jobs.

Frontend UIs

Kilo posted the full UI breakdown after running both models on ten prompts in Kilo Code CLI. Each task started in an empty directory, each model wrote a single self-contained index.html with Tailwind via CDN, and there were no follow-up prompts.

The result was weirdly convergent:

  • Same broad page anatomy across landing pages, dashboards, checkout, kanban, and docs.
  • Similar component choices, with differences mostly in palette, typography, and density.
  • Kimi K3 cost 29% of Fable 5 across the run.
  • Completed K3 runs averaged 9m 42s of agent time, versus 3m 50s for Fable 5, according to Kilo's writeup.

ReactBench gave a colder production-readiness check. Aiden Bai put K3 at #6, with a 33% score, a +10 point jump over K2.7, better than Opus 4.8 at over 2x lower cost, and still behind Fable and Sol.

DeepSWE

K3's coding story is strongest when the metric includes cost. Kolt Regaskes' broader launch summary cited DeepSWE at 69% for K3 at $4.65 per task, versus GPT-5.6 Sol at 73% and $8.39, and Fable 5 at 70% and $21.63.

The benchmark caveat lives in Kimi's own footnotes: K3 results were run at max reasoning effort, temperature 1.0, top-p 1.0, and across harnesses including KimiCode, Claude Code, and Codex depending on the benchmark. That makes harness choice part of the measurement, not background noise.

Office tasks

Composio tested K3 and Fable 5 on 14 agentic office tasks, including support-ticket spreadsheet sync, GitHub repo access audits, and CRM contact deduping. Both models scored 9 of 14, passing the same nine tasks and failing the same five.

The differences were operational:

  • Speed: Fable finished about 2.5x faster.
  • Calendar check: Fable 54 seconds, Kimi 200 seconds.
  • CRM cleanup: Fable 64 seconds, Kimi 274 seconds.
  • Token use: Kimi used about 40% fewer tokens across the run.
  • One task: Fable used 3.1M tokens, Kimi used 1.7M.

Composio later said the run was first-run for both models in one reply, used Claude Code for Fable and Pi for Kimi in another reply, and would be retested after a few days in a retest note.

Capacity

Kimi said demand over the first 48 hours pushed its GPUs close to current capacity. The company temporarily paused new subscriptions, kept existing subscribers active, and said it would split membership into Kimi Membership for web, app, and work, plus Kimi Code Membership for coding workflows.

The serving complaints lined up with that note. Cedric Chee measured OpenRouter first-party Kimi K3 at 19 tok/s and 7.90s p50 latency, while bridgemindai showed one OpenRouter provider at 16 tps and 11.22s latency.

Cline's self-hosting post used Kimi K2.6, not K3, but its chart put hard numbers on the deployment economics: $185K per month for pure API, $166K for floor-sized self-hosting plus API spillover, about $141K with autoscaling, and about $116K with aggressive tuning. The chart notes K3 economics will differ because of roughly 3x weight footprint.

Cline also shipped a launch-week access route: its ClinePass post advertised K3 on Cline CLI with about 5x discounted access, and a follow-up said the same OpenAI-compatible API could be used outside Cline.

Harness gotchas

Kimi's launch docs include a compatibility trap: K3 was trained in preserved thinking-history mode. If the harness does not pass back historical thinking content, or if a session switches to K3 midstream from another model, Kimi says generation quality may become highly unstable.

Two other launch details matter for agent builders:

  • K3 launched with max thinking effort by default, while low and high effort modes are scheduled for later, according to Kimi's tech blog.
  • K3 supports dynamically loaded tools, or deferred tools, which bigeagle_xd said can reduce initial context length when many tools or MCPs are available.
  • Newly loaded tools are appended to context rather than prepended, bigeagle_xd clarified in a follow-up.
  • Some harnesses expose preserve-thinking as a setting that is off by default, one reply noted after quoting Kimi's own limitation text.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 8 threads
TL;DR2 posts
AA-Briefcase1 post
Cost per task1 post
Frontend UIs1 post
DeepSWE1 post
Office tasks3 posts
Capacity4 posts
Harness gotchas2 posts
Share on X