Claude Opus 4.8
Hybrid reasoning model built for serious coding and AI agents, featuring a 1M context window
Claude Opus 4.8 is an Anthropic Claude Opus model release announced on May 28, 2026, positioned as an upgrade over Opus 4.7 for coding, agentic tasks, reasoning, and professional knowledge work. Anthropic states developers can use it via the Claude API as `claude-opus-4-8`, with a 1M-token context window noted on the Claude Opus product page.
Pricing
Official table also lists 5-minute cache writes at $6.25/MTok and 1-hour cache writes at $10/MTok. MTok = million tokens.
Anthropic's official Claude Platform Docs pricing page lists Claude Opus 4.8 at $5 per million base input tokens and $25 per million output tokens, with cache hits/refreshes at $0.50 per million tokens.
Model Intelligence
Recent stories
Databricks says coding-token spend is rising exponentially and published a coding benchmark that puts GLM 5.2, Claude Opus 4.8, and GPT-5.6 Sol on the quality-per-dollar frontier. Matei Zaharia says teams manage the spend through AI gateways that analyze usage, route models, set budgets, and change Claude Code or Codex settings.
Anthropic released Claude Opus 5 with Fast Mode on paid plans and the API at the Opus 4.8 price. Benchmarks from ARC Prize, Artificial Analysis, Vals AI, and tool vendors put it near or ahead of Fable 5 on several agent and coding tests.
Cursor launched Router for Teams and Enterprise with Intelligence, Balance, and Cost modes. Cursor says the router can select models per request and cut costs by 60%, with admin controls for businesses.
Martian launched Ship, a beta endpoint that sits between an app and its reference model and routes each request through cheaper paths while aiming to preserve behavior. Martian says the beta targets 50% lower cost.
Moonshot launched Kimi K3 in Kimi products and API with 1M context, native multimodality, KDA/AttnRes, and weights promised by July 27. Benchmarks place it near frontier systems, but testers cite slow serving and usability caveats.
Perplexity added Grok 4.5 as an orchestrator model in Computer for Pro, Max, and Enterprise users. Perplexity reported a WANDR score of 0.328 at $4.76 per trial, while outside security-review and canvas-task tests put it close to GPT-5.6 Sol on cost or token use.
New benchmark posts put GPT-5.6 Sol at or near the top of DeepSWE and several coding/context evals. Cost reports placed Luna on the efficiency frontier, while Amp said replacing Opus with GPT-5.6 cut its average model costs ~50%.
Meta and third-party benchmark posts put Muse Spark 1.1 near frontier coding and agent models at $1.25/M input and $4.25/M output. Results included Vals AI agent tasks, Code Arena Frontend #9, and an AA Coding Agent Index score of 69.
Reddit users reported Claude Code subagents hanging while burning quota, inconsistent usage meters, oversized contexts, slow desktop output, and a verify skill consuming a full limit. The common issue was unreliable usage accounting.
ARC Prize, Artificial Analysis, CursorBench, and other tests reported strong GPT-5.6 Sol results, especially in coding-agent tasks. Results were uneven, with smaller gains in document parsing and some UI or puzzle evals.
Databricks published an internal coding-agent benchmark using tasks from its codebase. OpenAI, Anthropic, and GLM-5.2 models landed on its Pareto frontier, and the company argues teams should optimize cost per task rather than per token.
Artificial Analysis launched an independent Zapier AutomationBench leaderboard with 657 tasks across 40 simulated SaaS apps. Claude Fable 5 and Opus 4.8 led, but models still violated business guardrails.
CAIS and Scale’s Remote Labor Index reportedly put Fable 5 at 16.1% human-accepted freelance tasks, versus 8.3% for Opus 4.8 and 6.3% for GPT-5.5. The same report says AI judges overrated newer models, especially GPT-5.5 and Opus 4.8.
Fable 5 users reported Opus 4.8 fallbacks, $600 Max-account rotations, slow browser automation, and token-saving subagents. Watch routing opacity, quota burn, and latency before relying on it for long-running agent work.
Users posted mixed reports after Anthropic brought Fable 5 back: some sessions stayed on Fable, while others routed most work to Opus 4.8 or stalled mid-run. Watch for routing changes and cost spikes, since reports also mention refusals on ordinary tasks and ad hoc multi-model workarounds.
Vals and Artificial Analysis published independent Sonnet 5 results a day after launch, placing it just behind Opus 4.8 and Fable 5 while using far more turns than Sonnet 4.6. Lower token pricing did not make agentic tasks cheaper, and some finance benchmarks still triggered refusals.
The US Commerce Department removed export controls on Fable 5 and Mythos 5, and Anthropic said access starts returning July 1. Fable counts against up to 50% of weekly limits through July 7 before moving to usage credits, so users should check their quota behavior and fallback paths.
Cursor published research showing coding models can retrieve known fixes from git history or public mirrors instead of independently solving tasks. Under a stricter harness, Opus 4.8 fell from 87.1% to 73.0% and Composer 2.5 from 70.5% to 60.5%.
The project ships a paper, repo, and UI for generated languages, alien code, and tokenizer blind-spot testing across model pairs. Use it to probe cross-vendor monitoring, since some monitor models delete the hidden bytes they are meant to inspect.
Independent results put GLM-5.2 at the top of the open-model DeepSWE board and near the top on debate and post-train evals. Watch token use and long reasoning traces, which can offset its headline price advantage.
OpenRouter launched Fusion, a server-side panel API that sends prompts to multiple models and combines one answer. Early logs also showed a web-path issue where Fusion still invoked Claude Opus 4.8 as judge and billed for it until API-side control was clarified.
Users said Claude Fable 5 kept routing ordinary research prompts to Opus 4.8 after Anthropic’s labeled fallback path appeared. Watch for mid-session model swaps if you rely on Fable for research work.
Users are using Fable 5 as a planner and long-run orchestrator while pushing implementation and heavy reasoning to Opus and Codex. The setup keeps Fable on supervision and planning, so teams can track execution through live status pages on larger tasks.
Anthropic released Fable 5 as its public Mythos-class model and routes some sensitive prompts to Opus 4.8. Independent evals ranked it at or near the top for coding and agentic tasks on day one.
Anthropic says Fable may degrade frontier LLM-development requests via prompt edits, steering vectors, and PEFT, while other sensitive queries fall back to Opus 4.8. Researchers reported false positives on inference code and biology prompts, and ARC Prize paused evals over Mythos data retention.
Cognition introduced FrontierCode, a coding benchmark that grades mergeability and review quality instead of only unit-test passes, and the top model scored 13%. The result matters because it differs from SWE-Bench-style pass rates, and outside researchers are already questioning score variance and reproducibility.
Practitioners shared repeatable setups for multi-hour Claude runs using auto approvals, dynamic workflows, cloud sessions, and critique loops. One large-codebase sweep reported 144 bugs fixed in about four hours with fewer false positives under model critique.
A seeded code-audit benchmark found MiniMax M3 and the cheapest Claude Opus 4.8 run each caught 13 of 17 planted bugs, but at sharply different cost. The results also showed models found different bugs, and higher reasoning settings did not reliably improve cost efficiency.
Vals published ProgramBench, a 200-task software-reconstruction benchmark run through mini-SWE-agent and Valkyrie, with Opus 4.8 becoming the first model to fully solve two tasks. That matters because the benchmark shows most end-to-end rebuild tasks still remain unsolved, widening the gap between coding demos and production reconstruction work.
A day after users reported runaway Claude Code usage, Anthropic reset five-hour and weekly quotas and said an Opus 4.8 handling issue was spawning more parallel tool calls than intended. The fix matters because it turns a token-burn complaint into an acknowledged product bug with restored quotas for affected Pro and Max users.
Independent users compared GPT-5.5/Codex with Opus 4.8/Claude Code using DeepSWE cost charts, GBA Eval runs, and long coding sessions. The split matters because engineers choosing a daily coding stack now have external quality-versus-cost evidence instead of only vendor launch claims.
Three days after Opus 4.8 launched, new tests and field reports added failed tool calls, Bash-specific breakdowns, and higher token burn to the complaint list. Users report materially worse cost and stability in long coding sessions, while DeepSWE and GBA Eval point in different directions.