Independent tests put Sonnet 5.5 at $7.60 per max-effort task
Sonnet 5.5 ranks second on Artificial Analysis' Intelligence Index but uses about 193,000 output tokens per task at max effort. A comparison of AA Index results puts its cost at $7.60 per task.

TL;DR
- Sonnet 5.5 reached 56 on the Artificial Analysis Intelligence Index, up from Sonnet 5’s 38, according to Artificial Analysis’s results.
- At max effort, the same test averaged $7.60 per task, compared with $5.98 for Opus 5.5 and $3.26 for GPT-6 Astra in haider1’s cost chart.
- The max-effort run generated about 193,000 output tokens per task, versus 119,000 for Opus 5.5, as haider1’s token chart shows.
- Cost moved differently on other workloads: ValsAI’s tests recorded a higher cost per test than Sonnet 5, while Box’s early results recorded fewer tokens and faster deliverables.
On Anthropic’s FrontierCode results, Max scored below Xhigh. In Simon Willison’s SVG test, Max exhausted its 128,000-token output allowance without producing the image. The developer guide sets different defaults for Claude Code and the API.
The $7.60 index task
Artificial Analysis measured a +18-point jump on its Intelligence Index, from 38 for Sonnet 5 to 56 for Sonnet 5.5 at Max, two points behind Opus 5.5. Its independent report puts the new model’s average cost at $7.60 per index task, above Opus 5.5’s $5.98 despite Sonnet’s lower token rates.
Anthropic’s price sheet keeps Sonnet 5’s rates of $2 per million input tokens and $10 per million output tokens, half Opus 5.5’s $4 and $20. Anthropic reports up to 30% lower cost per task than Sonnet 5 in its own testing; the Artificial Analysis figure is for its max-effort index workload.
The 193,000-token output
Artificial Analysis counted roughly 193,000 output tokens per index task at Max, about 62% more than Opus 5.5’s 119,000 and roughly seven times GPT-6 Astra’s 27,000. These are totals across tasks, while the model’s 128,000-token output limit applies to a single response.
Effort settings and the cost curve
Sonnet 5.5 has five tested settings, low, medium, high, xhigh and max, in Artificial Analysis’s comparison post. A separate cost chart puts Xhigh at $2.74 per index task against Max at $7.60. Anthropic sets Medium as the Claude Code and apps default, and High as the Claude Platform default, according to its launch post.
More effort also failed to improve one coding score: on Anthropic’s FrontierCode 1.1 main set, Xhigh scored 52.1% and Max 46.2%. Anthropic’s footnote says Max more often invoked a multi-subagent code-review skill; Cognition found two cases that timed out or made edits beyond the task’s scope.
Vals Index: accuracy and cost rose together
On ValsAI’s index, Sonnet 5.5 moved from Sonnet 5’s 59.61% accuracy to 69.22%, +9.61 points, while cost per test rose from $11.74 to $20.80. Its 18-benchmark comparison showed 16 gains and two declines:
- CyberBench: 62% → 60%, approximately −2.3 points on ValsAI’s unrounded scores.
- Harvey Legal Agent: 5% → 3%, approximately −2.1 points on ValsAI’s unrounded scores.
ValsAI’s tests and Artificial Analysis’s index measure different task mixes; their $20.80 and $7.60 figures are averages over their respective tests.
Box’s shorter enterprise runs
Box reported a different token pattern in early access testing with Box Agent: overall accuracy rose from 61% to 65%, +4 points, while finished deliverables arrived roughly 2.4 times faster using 12% fewer tokens. Its financial-services subset improved from 63% to 81%, +18 points, according to Box’s results.
Pre-release test caveat
Artificial Analysis ran its evaluations on a pre-release deployment with a structured-output bug that Anthropic says it fixed for the public release. The evaluator plans to rerun affected tests; it also observed fallback to Sonnet 5 on about 0.1% of index tasks, primarily in Terminal-Bench 4.0.