Skip to content
AI Primer
update

Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

5 min read
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing
Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing

TL;DR

  • Kimi K3’s harness tax finally has a dollar figure: success stayed in a 71 to 79% band across three harnesses while median token use ranged from 61K to 340K, as thursdai_pod’s summary put it.
  • Median per-task cost swung from $0.22 in Kimi Code to $2.00 in Claude Code, according to thursdai_pod’s cost breakdown.
  • Cline used Kimi K3 to improve its own harness, moving Terminal-Bench 2.1 from 77.5% to 88.8% and run cost from $79 to $49.8 in 17 hours, per Cline’s RSI post.
  • The rollout is already in tools: Cline’s CLI update shipped npm i -g cline, Kimi_Moonshot named Together as a day-zero partner, and OpenCode added Kimi K3 to OpenCode Zen.
  • Harness compatibility is still messy: HamelHusain called K3 in Codex Desktop a “big mistake,” while bigeagle_xd pointed to Codex’s Responses API requirement.

Kimi’s Hugging Face model card describes a 2.8T-parameter open-weight multimodal model with a 1M-token context window, and the technical report says KDA, Attention Residuals, and Stable LatentMoE drove a 2.5x scaling-efficiency gain over K2. OpenRouter’s Kimi K3 page already lists multi-provider routing for the same model at $3/M input and $15/M output. Cline’s blog post adds the agent nerd candy: a single recursive self-improvement run that turned harness patches into a Terminal-Bench score jump.

The 28-task harness spread

Composio ran Kimi K3 through Kimi Code, Hermes, and Claude Code on the same 28 tasks, according to thursdai_pod’s summary. The token counts came from thursdai_pod’s median-token post, and the success and latency counts came from the outcome post.

| Harness | Solved | Median tokens | Median time |
| --- | ---: | ---: | ---: |
| Kimi Code | 22/28 | 61K | 297s |
| Hermes | 21/28 | 67K | 179s |
| Claude Code | 20/28 | 340K | 348s |

Claude Code used 5.6x more median tokens than Kimi Code while solving two fewer tasks. Hermes was the fastest run, with token use close to Kimi Code.

Per-task dollars

Prices widened the gap: thursdai_pod’s cost post put Kimi Code at $0.22/task, Hermes at $0.28/task, and Claude Code at $2.00/task. That is a 9.1x spread for a 7.1-point success-rate spread.

Composio’s separate 14-task agentic benchmark said K3 matched Claude Fable 5 task for task, with 9 passes and 5 failures, and did better on the hardest task, 11 of 13 checks versus Fable’s 9 of 13 in Composio’s comparison. The same post said roughly 95% of token spend on agentic tasks comes from input tokens.

Composio called the price gap “pretty dramatic” in a reply. When asked about Kimi taking longer in the comparison, Composio said the delay was probably hosting-related in another reply.

Cline’s recursive harness patch

Cline’s experiment is the cleaner harness story: one 17-hour run took Kimi K3 in the Cline harness from 77.5% to 88.8% on Terminal-Bench 2.1 and cut run cost from $79 to $49.8, according to Cline’s RSI post. The attached chart labels the after state as “one prompt, zero benchmark-specific fixes.”

Cline’s full-run cost chart put the same $49.8 Kimi K3 run against $400 for GPT-5.6 Terra and $552 for Claude Fable 5 in its Terminal-Bench comparison. Cline also said the open-source harness can be forked and rerun with another model in the follow-up, while the blog post includes the writeup.

Open weights, hosted routing

Moonshot teased Kimi K3 as “open weights, coming soon” in Kimi_Moonshot’s preview, and bridgemindai captured the provider race version of the release: 2.8T parameters on Hugging Face, with fast inference providers getting a copy.

Day-zero access spread across agent surfaces fast:

  • Cline: Cline’s CLI update says npm i -g cline exposes Kimi’s self-improvements, and Cline’s ClinePass reply describes a new-user promo with about $60 of Kimi K3 API usage per month.
  • Together: Kimi_Moonshot thanked Together AI as a day-zero launch partner for K3 access optimized for coding agents and production workloads.
  • OpenCode: OpenCode added Kimi K3 to OpenCode Zen.
  • Morph: Morph made Kimi K3 available directly, through OpenRouter, and through Vercel Gateway.

Closed-harness friction

The awkward edge case is closed harnesses. HamelHusain’s test called K3 in Codex Desktop a “big mistake,” and his reply said it “gets really weird and sends json messages.”

A protocol mismatch may explain part of it: bigeagle_xd said Codex needs the Responses API and Kimi does not provide it. In a follow-up, HamelHusain said OpenCode and pi were fine, while customized harness behavior can get in the way.

OpenBench is turning that friction into something measurable. mattlam_ linked an easier runner, and mattlam_’s caveat said Cursor data was self-reported while Codex repeatedly showed high token usage.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR3 posts
The 28-task harness spread1 post
Per-task dollars2 posts
Cline’s recursive harness patch1 post
Open weights, hosted routing5 posts
Closed-harness friction4 posts
Share on X