Composio benchmarks Kimi K3 harnesses with $0.22–$2 per-task cost swing
Composio found Kimi K3 success stayed near 71–79% across three harnesses while median token use ranged from 61K to 340K and cost from $0.22 to $2 per task. Cline also shipped a Kimi K3 CLI update.

TL;DR
- Kimi K3’s harness tax finally has a dollar figure: success stayed in a 71 to 79% band across three harnesses while median token use ranged from 61K to 340K, as thursdai_pod’s summary put it.
- Median per-task cost swung from $0.22 in Kimi Code to $2.00 in Claude Code, according to thursdai_pod’s cost breakdown.
- Cline used Kimi K3 to improve its own harness, moving Terminal-Bench 2.1 from 77.5% to 88.8% and run cost from $79 to $49.8 in 17 hours, per Cline’s RSI post.
- The rollout is already in tools: Cline’s CLI update shipped
npm i -g cline, Kimi_Moonshot named Together as a day-zero partner, and OpenCode added Kimi K3 to OpenCode Zen. - Harness compatibility is still messy: HamelHusain called K3 in Codex Desktop a “big mistake,” while bigeagle_xd pointed to Codex’s Responses API requirement.
Kimi’s Hugging Face model card describes a 2.8T-parameter open-weight multimodal model with a 1M-token context window, and the technical report says KDA, Attention Residuals, and Stable LatentMoE drove a 2.5x scaling-efficiency gain over K2. OpenRouter’s Kimi K3 page already lists multi-provider routing for the same model at $3/M input and $15/M output. Cline’s blog post adds the agent nerd candy: a single recursive self-improvement run that turned harness patches into a Terminal-Bench score jump.
The 28-task harness spread
Composio ran Kimi K3 through Kimi Code, Hermes, and Claude Code on the same 28 tasks, according to thursdai_pod’s summary. The token counts came from thursdai_pod’s median-token post, and the success and latency counts came from the outcome post.
| Harness | Solved | Median tokens | Median time |
| --- | ---: | ---: | ---: |
| Kimi Code | 22/28 | 61K | 297s |
| Hermes | 21/28 | 67K | 179s |
| Claude Code | 20/28 | 340K | 348s |
Claude Code used 5.6x more median tokens than Kimi Code while solving two fewer tasks. Hermes was the fastest run, with token use close to Kimi Code.
Per-task dollars
Prices widened the gap: thursdai_pod’s cost post put Kimi Code at $0.22/task, Hermes at $0.28/task, and Claude Code at $2.00/task. That is a 9.1x spread for a 7.1-point success-rate spread.
Composio’s separate 14-task agentic benchmark said K3 matched Claude Fable 5 task for task, with 9 passes and 5 failures, and did better on the hardest task, 11 of 13 checks versus Fable’s 9 of 13 in Composio’s comparison. The same post said roughly 95% of token spend on agentic tasks comes from input tokens.
Composio called the price gap “pretty dramatic” in a reply. When asked about Kimi taking longer in the comparison, Composio said the delay was probably hosting-related in another reply.
Cline’s recursive harness patch
Cline’s experiment is the cleaner harness story: one 17-hour run took Kimi K3 in the Cline harness from 77.5% to 88.8% on Terminal-Bench 2.1 and cut run cost from $79 to $49.8, according to Cline’s RSI post. The attached chart labels the after state as “one prompt, zero benchmark-specific fixes.”
Cline’s full-run cost chart put the same $49.8 Kimi K3 run against $400 for GPT-5.6 Terra and $552 for Claude Fable 5 in its Terminal-Bench comparison. Cline also said the open-source harness can be forked and rerun with another model in the follow-up, while the blog post includes the writeup.
Open weights, hosted routing
Moonshot teased Kimi K3 as “open weights, coming soon” in Kimi_Moonshot’s preview, and bridgemindai captured the provider race version of the release: 2.8T parameters on Hugging Face, with fast inference providers getting a copy.
Day-zero access spread across agent surfaces fast:
- Cline: Cline’s CLI update says
npm i -g clineexposes Kimi’s self-improvements, and Cline’s ClinePass reply describes a new-user promo with about $60 of Kimi K3 API usage per month. - Together: Kimi_Moonshot thanked Together AI as a day-zero launch partner for K3 access optimized for coding agents and production workloads.
- OpenCode: OpenCode added Kimi K3 to OpenCode Zen.
- Morph: Morph made Kimi K3 available directly, through OpenRouter, and through Vercel Gateway.
Closed-harness friction
The awkward edge case is closed harnesses. HamelHusain’s test called K3 in Codex Desktop a “big mistake,” and his reply said it “gets really weird and sends json messages.”
A protocol mismatch may explain part of it: bigeagle_xd said Codex needs the Responses API and Kimi does not provide it. In a follow-up, HamelHusain said OpenCode and pi were fine, while customized harness behavior can get in the way.
OpenBench is turning that friction into something measurable. mattlam_ linked an easier runner, and mattlam_’s caveat said Cursor data was self-reported while Codex repeatedly showed high token usage.