Fireworks says Ember-1 uses 71% fewer reasoning tokens on coding tasks
Fireworks says Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens than K3 in a live coding-traffic test. The models achieved the same success rate.

TL;DR
- Fireworks post-trained Kimi K3 into Ember-1 to spend fewer tokens on repeated reasoning, with roughly 40% fewer tokens at comparable benchmark quality, according to the launch summary.
- In a live coding-traffic A/B test, Fireworks reported 71% fewer reasoning tokens and 39% fewer total tokens at the same success rate, as quoted in cline's post.
- The numbers need a footnote: the 39% claim differs from the 34.5% total-token reduction in Fireworks' published customer table.
- Ember-1 is available in Cline and on Vercel AI Gateway, according to cline's availability post and the AI Gateway announcement.
Simply lowering K3's reasoning effort lost too much quality, Fireworks says, so it trained a new variant. The same benchmark table contains two coding-score declines alongside the token savings. Its model card lists a 1.04-million-token context window.
The live coding-traffic test
Fireworks' headline A/B result comes from production coding traffic, rather than a synthetic prompt set. cline's account of the test gives the reductions as 71% for reasoning tokens and 39% for total tokens, with no change in success rate.
The September 23 launch post describes A/B tests with two customers and approximately 35% fewer tokens per task at comparable quality. Its published table reports 34.5% fewer total tokens and a score of 0.750 for K3 versus 0.753 for Ember-1. Fireworks does not explain why that table's total-token figure differs from the 39% in the later post, or give a task count for the customer tests. One customer moved Ember-1 into production after its test, the company says.
Why shorter traces compound in agents
Reasoning can account for more than 90% of a model's generated tokens, according to Fireworks' technical account. When an agent sends earlier turns back with each new request, previous reasoning is read and billed again; Fireworks describes context growth across such turns as roughly quadratic.
The training targeted repetitive thought while preserving reflection that changes an answer, including revisiting an assumption after tool feedback. The distinction is especially consequential in coding loops, where a long trace from an early step can travel through every later call.
Post-training rather than an effort setting
Fireworks says K3's lower reasoning-effort settings cost too much accuracy. The team instead ran more than 50 training experiments and 200 evaluations on Fireworks Serverless Training, using task feedback to shape reasoning across standalone problems and extended interactions, according to its training account.
The task collection covered:
- Mathematics
- Coding
- Instruction following
- Conversation
- Search
- Tool use
- Software engineering
That breadth is part of Fireworks' claim that the savings transfer beyond one coding benchmark; its launch post reports shorter reasoning across seven benchmarks and two customers' production traffic. cline's follow-up called Ember-1 the first model from the new Fireworks Research team.
Five benchmark scores, including two dips
The Fireworks benchmark table compares Ember-1 with K3 at maximum reasoning effort. These are company-reported comparisons, including the Ember-1 result shown on the Terminal-Bench chart attached to the launch post.
- Terminal Bench 2.1: 80.9% → 82.0%, +1.1 points (89 tasks).
- SWE-bench Verified: 93.2% → 92.2%, −1.0 point (500 tasks).
- SWE-Interact: 21.3% → 20.0%, −1.3 points (75 tasks).
- DeepSWE 1.1: 66.4% → 75.2%, +8.8 points (113 tasks).
- τ-2 Bench Airline: 64% → 66%, +2 points (50 tasks).
Fireworks calculated its cost comparisons with public K3 rates: $3 per million uncached input tokens, $0.30 per million cached input tokens, and $15 per million output tokens. omarsar0's cost-versus-score chart makes the efficiency case visually, while the per-benchmark results show where quality moved in either direction.
Serverless access and the two-week preview
The Fireworks model card lists the API path as accounts/fireworks/models/ember-1, with serverless pricing of $3 per million input tokens, $0.30 per million cached input tokens, and $15 per million output tokens. Function calling and on-demand deployment are also listed. Those are the same rates used in Fireworks' K3 benchmark cost calculations; the claimed savings come from fewer tokens consumed, rather than a lower per-token rate.
Fireworks calls this a Research Preview with two weeks of serverless access and says permanent availability depends on community demand in its launch post. cline's post says the model is available in Cline, including its Desktop app. Vercel lists it as fireworks/ember-1 on AI Gateway with image input, tool use, and Zero Data Retention, per its announcement and its availability note.