Anthropic releases Claude Sonnet 5.5 at unchanged token prices
Sonnet 5.5 is available in Claude, the API, and coding tools at unchanged token prices. Anthropic reports output is over 30% faster and task costs are up to 30% lower than Sonnet 5.

TL;DR
- Claude Sonnet 5.5 shipped across Claude and the API; Anthropic's announcement reports over 30% faster output generation and up to 30% lower cost per task than Sonnet 5.
- Token prices remain $2 per million input and $10 per million output, with $0.20 cache reads, according to Claude's developer update.
- Agentic coding made the biggest jump in Anthropic's table: Terminal-Bench 4.0 rose from 10.3% to 70.6%, +60.3 points, as the launch-table screenshot shows.
- Max effort can erase the savings: Artificial Analysis' independent evaluation measured $7.60 per task, roughly 50% more than Sonnet 5 on its Intelligence Index.
- The API migration changes tool selection and thinking behavior; Claude's building guide documents five breaking changes beyond the model ID swap.
Anthropic's FrontierCode footnote explains a rare benchmark dip at Max effort: extra review subagents sometimes timed out or changed too much. The migration guide says forced tool selection now returns a 400, while Artificial Analysis measured 193,000 output tokens per task at Max. Simon Willison's SVG test consumed the full 128,000-token output allowance without producing the requested pelican.
What shipped
- Model and access: Sonnet 5.5 is the second Claude 5.5 model after Opus; Haiku 5.5 is due in the coming weeks, per Anthropic's rollout post. Sonnet 5.5 powers Claude's free chat tier, according to Simon Willison's free-tier report; his clarification says Claude Code still requires a paid account.
- API and clouds: The model slug is
claude-sonnet-5-5on the Claude API, Claude Platform on AWS, Google Cloud, and Microsoft Foundry; Bedrock usesanthropic.claude-sonnet-5-5, as the developer guide specifies. - Price: Input/output rates stay at $2/$10 per million tokens; 5-minute cache writes cost $2.50 per million, and cache reads cost $0.20, per Claude's pricing post. Anthropic attributes its up-to-30% task savings to fewer tokens, while its speed-and-cost post puts output generation more than 30% faster.
- Defaults: Claude apps and Claude Code start at Medium effort, while the Platform defaults to High, according to Claude's developer note. the Claude Code 2.1.284 changelog says Sonnet 5.5 replaced Sonnet 5 as the default Sonnet selection, not the overall Opus default.
- Usage reset: Pro, Max, and Team accounts received a reset that can be applied before October 22, according to Claude's rollout note.
Benchmarks that moved
The launch table mixes model effort settings; its Opus Terminal-Bench comparator is Opus's best Xhigh run. Scores below compare Sonnet 5 with Sonnet 5.5 within each source's own evaluation.
First-party
- Terminal-Bench 4.0: 10.3% → 70.6%, +60.3 points, per the launch-table screenshot.
- CursorBench 4.0: 34.1% → 55.5%, +21.4 points, per the benchmark table.
- FrontierCode 1.1 Main at the reported best settings: 42.4% → 52.1% at Xhigh, +9.7 points, per Anthropic's table.
- GDPval-AA v2.1: 1,449 → 1,844 Elo, +395 Elo, per the launch numbers.
- OSWorld 2.1 partial: 57.0% → 80.1%, +23.1 points, per the comparison table.
- Chartography without tools: 15.6% → 61.6%, +46.0 points, per the published chart.
Third-party evaluators
- Artificial Analysis Intelligence Index at Max: 38 → 56, +18 points, per Artificial Analysis' release analysis.
- Artificial Analysis' Terminal-Bench 4.0: 14.1% → 63.6%, +49.5 points, per its independent runs. Its absolute score differs from Anthropic's 70.6% result; the two organizations report separate runs.
- Vals Index accuracy: 59.61% → 69.22%, +9.61 points, per Vals AI's evaluation thread.
Customer-reported
- Box's complex-work eval: 61% → 65%, +4 points, per Box's early-access results; its financial-services subset rose 63% → 81%, +18 points, in the same tests.
- Cognition's FrontierCode 1.1 run in Devin: 56.2% → 64.4%, +8.2 points, per Cognition's Devin launch thread. This is Cognition's agent evaluation, rather than the score in Anthropic's launch table.
Where it regressed
More effort sometimes produced more work and a worse result. Anthropic's FrontierCode footnote reports Sonnet 5.5 at Xhigh scoring 52.1% versus 46.2% at Max, −5.9 points; in two cases Cognition examined, extra code-review subagents caused a timeout or out-of-scope edits.
- Vals measured CyberBench at 62% → 60%, −2 points, and Harvey Legal Agent at 5% → 3%, −2 points, in its 18-benchmark thread. The other 16 benchmarks rose.
- Artificial Analysis found factual accuracy of 54% for Sonnet 5.5 versus 66% for Opus 5.5 on AA-Omniscience, −12 points relative to Opus, according to its evaluation. It also reports lower hallucination rates for Sonnet on that test, 47% versus 59%.
- Anthropic says tasks can cost up to 30% less than Sonnet 5 in its own testing, while Artificial Analysis' Max-effort runs cost $7.60 per Intelligence Index task, about 50% more than Sonnet 5 and above Opus 5.5's $5.98. The measurements cover different workloads and effort settings; flat token rates do not fix the task bill.
At Max, Artificial Analysis counted about 193,000 output tokens per task, versus 119,000 for Opus 5.5. Vals also reported occasional 1-million-token context exhaustion and sandbox crashes from too many processes. Simon Willison's hands-on test hit the 128,000-token output limit without an SVG at Max; Xhigh produced one in 41 seconds for 5.74 cents.
Artificial Analysis tested a pre-release deployment with a structured-outputs bug. Anthropic says it fixed the bug for public release, and Artificial Analysis said it would rerun affected evaluations.
Under the hood
The developer model sheet lists a native 1-million-token context window, 128,000 maximum output tokens, five effort levels (low, medium, high, xhigh, max), and the same tokenizer as Sonnet 5. Message Batches can reach 300,000 output tokens with a beta header; the minimum cacheable prompt drops from 1,024 to 512 tokens. A 2000×1500 image costs roughly 2.5 times as many input tokens as on Sonnet 4.6, Sonnet 4.5, or Haiku 4.5 because Sonnet 5.5 uses the high-resolution image tier.
The Sonnet 5.5 migration guide spells out the Messages API changes behind Claude's short developer post:
- Thinking: Adaptive thinking remains on when the field is omitted. Sonnet 5's
thinking.type: "disabled"now returns a 400;between_toolsremoves up-front thinking at Low, Medium, or High, but returns a 400 at Xhigh or Max. Thinking counts toward billed output andmax_tokens. - Tool choice: Forced
tool_choice: "tool"or"any"returns a 400.autoplusstrict: trueconstrains tool inputs without requiring a tool call; Bedrock lacks strict-tool support for this model. - Conversation state: Responses can start with a
thinkingblock, so clients readingcontent[0].textbreak. Thinking blocks must travel unchanged through tool loops; progress between tool calls can arrive in those blocks rather than text. - Computer use and advisors: The Claude API and Google Cloud require
computer_toolset_20260801instead ofcomputer_20251124; Bedrock still accepts the older tool. Sonnet 5.5 also rejects Opus 4.8, Opus 4.7, and Sonnet 5 as advisors, and returns accepted advisors' advice encrypted. - Preserved thinking: Sonnet 5.5 can read Sonnet 5's thinking blocks, but older models cannot read Sonnet 5.5's. Anthropic says thinking remains bound to its originating organization; switching accounts mid-session makes Claude reread the session and generate fresh thinking, per the developer announcement.
Sonnet 5.5 is the first Sonnet shipped with cyber safeguards and fallback routing for higher-risk requests, according to Anthropic's safety note. The release details say flagged cyber requests visibly fall back to Sonnet 5, while routine software development is unaffected; Artificial Analysis observed fallback on roughly 0.1% of its Intelligence Index tasks.
Vibe Check
After a day of use, kunchenguid found Opus better at planning ambiguous product features and Sonnet faster at implementation and orchestration in their first-day report. Their working split was Opus for plans, Sonnet for implementation, and the forthcoming Haiku for fixes.
- A small internal-library update went from prompt to PR in about 20 seconds, according to rishdotblog's first run.
- Sonnet handled revisions better than Astra in Dan Shipper's work tests, though Astra remained his favorite overall writing model; quick coding and design iteration favored Sonnet for his colleague, per Shipper's hands-on assessment.
- A photo-to-code painting test had the model write a Python brush engine, render its work, inspect it, and revise it; Lance Martin posted the Sonnet 5 comparison.
- An interactive three.js plasma globe felt close to Opus at half the token price to one builder who made it. Another side-by-side voxel test took Sonnet 5.5 at High 49 minutes and $9.23, versus 22 minutes and $4.85 for Sonnet 5 at High, despite the richer output.
Where it shows up
Cursor reported that Sonnet 5.5 performs on par with Opus on many of its tasks in its launch post. Other day-one surfaces span coding agents, gateways, and browser tools:
- Coding agents: Devin Desktop and CLI added Sonnet 5.5, per Cognition's rollout thread; Factory's first tests say High effort checked the actual requirement rather than merely the nearest failing test. Warp's announcement covers both Warp and its Agent CLI.
- Gateways: OpenRouter's launch lists a 1-million-token context window at $2/$10 per million input/output tokens, while Vercel's AI Gateway release exposes
anthropic/claude-sonnet-5.5. - App builders: v0's release added the model, and Julius's rollout reports stronger coding and clearer writing inside its analysis, presentation, and website workflows.
- Browser agents: Hyperbrowser's launch put Sonnet 5.5 in its Agents Playground for navigation, clicks, and typing.