Agent Arena ranks Claude Opus 5 Max No. 2 across 7,000+ agent sessions
Agent Arena says Claude Opus 5 Max placed second and Opus 5 High placed third across more than 7,000 real agentic sessions. Practitioner reports remain mixed across harnesses and prompting styles.

TL;DR
- Opus 5 Max landed at No. 2 with 11.88% net improvement over 7,253 sessions, and Opus 5 High landed at No. 3 with 11.73%, according to Agent Arena's post.
- Pricing stayed at $5 per million input tokens and $25 per million output tokens, Agent Arena's pricing note says.
- The cost curve is tighter than the rank suggests: Agent Arena's chart puts Opus 5 High near Fable 5 High on net improvement at roughly the same median cost per task.
- The migration gotcha is context bloat: omarsar0's later pass says Opus 5 worked better after cleaning CLAUDE.MD, skills, and tool descriptions, while rohanpaul_ai's clip quotes Boris Cherny telling Claude Code users to try deleting old scaffolding.
- Independent evals split hard: datacurve's DeepSWE result put Opus 5 at 74% for long-horizon coding, while composio's summary found GPT-5.6 Sol passed more agentic tasks with fewer tokens, less time, and lower cost.
The live Agent Arena leaderboard is based on 1,412,751 sessions and 44 models, and its methodology treats agent evaluation as a causal tracing problem rather than pairwise voting. Anthropic's launch post frames Opus 5 as near-Fable intelligence at half the price, but Claude Code's memory docs add a key caveat: CLAUDE.md and auto memory are context, not enforced configuration. The strange corners include a live claim that Opus 5 becomes much better when old skills are removed omarsar0's later pass, and a prompt-game thread where strings like “Fable” or “Dario and Amanda” allegedly push the model into different simulated voices fabianstelzer's Promptemons post.
Agent Arena ranking
The live Agent Arena leaderboard listed these top rows on July 28:
- Claude Fable 5 High: rank 1, 12.58% ± 2.19% net improvement, 23,807 sessions.
- Claude Opus 5 Max: rank 2, 11.88% ± 2.81% net improvement, 7,253 sessions.
- Claude Opus 5 High: rank 3, 11.73% ± 1.66% net improvement, 11,114 sessions.
- GPT-5.6 Sol xHigh: rank 4, 10.02% ± 1.63% net improvement, 16,966 sessions.
Agent Arena says the tasks are real long-horizon agent sessions with web search, filesystem, and terminal tools in the announcement. Its methodology post says causal tracing randomizes components in a multi-intervention trial, then estimates each component's effect as “net improvement.”
Signal split
Opus 5 Max led two of Agent Arena's visible signals, while Fable 5 and GPT-5.5 led others.
- Confirmed Success: Opus 5 Max, 17.56% ± 4.69%.
- Praise vs Complaint: Opus 5 Max, 25.63% ± 9.98%.
- Steerability: Fable 5 High, 12.93% ± 4.49%.
- Bash Recovery: GPT-5.5 xHigh, 14.14% ± 1.19%.
- Tool Hallucination: Kimi K3 Max, 1.32% ± 0.19%, listed as the lowest hallucination signal on the page.
The broader Arena picture was stronger on visual and text tasks. Agent Arena's follow-up put Opus 5 Max at rank 1 in Frontend Code Arena and rank 1 in Text Arena with factuality on, while Opus 5 High ranked third in Frontend Code and second in Text.
Cost curve
The test-time scaling chart makes the practical tradeoff visible:
- Opus 5 Max: about +11.9% net improvement at about $2.50 median cost per task.
- Opus 5 High: about +11.7% at about $1.50.
- Opus 5 Medium: about +10.2% at about $1.00.
- GPT-5.6 Sol xHigh: about +10.0% at about $1.00.
- Fable 5 High: about +12.6% at about $1.50.
The pricing line stayed flat versus recent Opus pricing, Agent Arena's pricing note says: $5 per million input tokens and $25 per million output tokens.
Context cleanup
The strongest hands-on reversal came from omarsar0. After initially saying Opus 5 broke workflows and ignored skills omarsar0's first reaction, he later wrote that it was “trained to be more agentic than anything I've used” and that his old context stack was the problem omarsar0's later pass.
His working changes were concrete:
- Keep persistent system prompts and CLAUDE.MD lightweight.
- Remove memories and tool descriptions from persistent context.
- Split situational context from persistent context.
- Use progressive disclosure through commands, skills, and links.
- Deduplicate MCP tool descriptions from the system prompt.
- Strip conflicting reliability instructions that older workflows accumulated.
Boris Cherny, identified in rohanpaul_ai's clip as the Claude Code creator, gave the harsher version at Y Combinator Startup School 2026: every six months, delete CLAUDE.md, skills, and hooks, then see what the model does.
Hands-on reports
The user reports clustered around the same failure mode: Opus 5 can look thorough while doing too much.
- It “goes too far,” treats comments as P0 issues, and writes good code alongside dumb mistakes, theo's follow-up said.
- It writes liked code, asks good questions, and feels fast, but overcomplicates work, misses the point, does unrelated extras, and has trouble stopping, davis7's notes said.
- It caused “boneheaded bugs and performance regressions” across projects, according to doodlestein's reply, after doodlestein's recovery post described using Fable agents to repair the damage.
- It was “blustering and often wrong” as a collaborator, Steve_Yegge's post said.
- It underperformed an open-weight harness on very long tasks, according to Tim_Dettmers's post, and Tim_Dettmers's harness reply said the harness is optimized for tasks above 10 million tokens.
datacurve found a behavioral difference that matches the vibe reports: Opus 5 commented on only 11% of tool calls, versus 65% for Fable, but Opus always gave a summary after working datacurve's tool-call note.
Benchmark split
The public evals do not line up cleanly.
- DeepSWE: Opus 5 scored 74%, which datacurve's result called the best long-horizon coding model it had seen; datacurve's cost note said Opus 5 cost 45% less per task than Fable 5.
- ReactBench: Opus 5 scored 49%, and aidenybai's post called it the best Claude model on ReactBench and more than 2x cheaper than Fable 5.
- Frontier-Bench: fastinoAI's Pioneer post said Opus 5 more than doubled Opus 4.8, 43.3% versus 21.1%, at lower cost per task.
- ProgramBench: scaling01's ProgramBench post put Opus 5 at 41.50% accuracy, ahead of Fable 5 at 33.00% and GPT-5.6 Sol at 23.00%.
- LisanBench: scaling01's LisanBench post said Opus 5 High was No. 1 overall but used many more tokens than Opus 4.8 High.
- PostTrainBench v1.1: karinanguyen's update put Fable 5 Max first at 41.8%, GPT-5.6 Sol Max second at 36.2%, and Opus 5 at 34.1%.
- Composio's 23-task agent benchmark: composio's summary said GPT-5.6 Sol passed 96% versus 87% for Opus 5, used 377K versus 699K tokens per task, and cost about $43 versus $80 for the suite.
Composio's failure examples were not small style misses. composio's failure breakdown said Opus labeled only 1 of 28 Gmail messages on one task, then logged chargeback orders it was told to skip and emailed those customers on a refund-ledger task.
3D harnesses
The most shareable Opus 5 demos were not repo-maintenance tasks. They were games, worlds, and simulations.
- kimmonismus's post said Opus 5 ran for 24 hours and built an entire game from scratch without external assets or code.
- omarsar0's flight-sim post showed a Three.js flight simulator and credited a judge-executor harness.
- omarsar0's harness note said the harness dynamically spins out a subagent to critique graphics and playability.
- nptacek's read called Opus 5 generally weaker on many tasks but better at 3D spatial reasoning than any model he had seen.
- skirano's submarine post showed a one-shot submarine game with custom assets, models, textures, and music.
mattshumer_ gave the technique a name: Gauntlet Loop. mattshumer_'s guide link describes a lead agent splitting a goal into parts, assigning each part a builder and a fresh critic, then iterating each component against a concrete quality bar.
Promptemons and memory oddities
fabianstelzer claimed certain addressees push Opus 5 into a simulated “base model” mode. In his examples, “Dario and Amanda” produced corporate sci-fi email prose, “Fable” produced random multilingual prose, and “Sam” or “Ilya” did not trigger the same behavior fabianstelzer's Promptemons post.
Memory behavior had its own edge case. nptacek's screenshot showed Opus dismissing a nostalgic tangent as “insufficiently novel for documentation,” then recalling and updating memory anyway nptacek's memory screenshot; nptacek's reply said it did record the memory in the end.