Skip to content
AI Primer
breaking

Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost

Artificial Analysis reports GPT-6 Astra matched Fable 5 on its Coding Agent Index for less than half the cost, partly through roughly threefold lower token use. Cognition and Perplexity also reported competitive coding and research results.

5 min read
Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost
Artificial Analysis reports GPT-6 Astra matches Fable 5 coding at under half the cost

TL;DR

Codex now keeps notes across context windows and searches earlier tool output, per OpenAI’s launch post. The Devin announcement measures whether patches are mergeable, while Simon Willison’s independent note flags that Astra’s 99.9% ARC-AGI-3 result used a $19,000 stateful provider-adapter harness, versus 62.7% with the default harness.

$4.72 coding task

The Artificial Analysis breakdown compares agents inside named products, not bare model calls: Astra ran in Codex, Fable in Claude Code, and Spark in Muse Code. That harness label belongs beside the score.

  • Score: At max effort, Astra reached 67. Fable 5.1 led the index at 70, while Fable 5, Opus 5, and Spark 1.3 sat around Astra’s mark, per Artificial Analysis's benchmark report.
  • Cost: Artificial Analysis's comparison shows Astra max at $4.72 per task, about the same as Sol max at $5.00. Artificial Analysis says the same score cost less than half as much as Fable 5.
  • Tokens: Astra used about one-third of Sol’s total tokens and one-fifth of Opus 5’s in the tested coding configurations, according to Artificial Analysis's benchmark report.

$10/$50 trade

The unit-price change is stark: OpenAI moved from Sol’s $4 per million input and $20 per million output tokens to $10 and $50. Cache reads retain a 90% discount and cache writes a 25% premium, according to Artificial Analysis’s pricing analysis.

On Artificial Analysis’s broader Intelligence Index, Astra max scored 61, equal to Sol and five points behind Fable 5.1. Its roughly 10% output-token reduction left it 75% more expensive per task than Sol at max effort, as Artificial Analysis's chart shows.

The model did improve on individual components: its AA-Omniscience hallucination rate fell from 92% to 51% while accuracy gained four points. AA also recorded roughly an 80-Elo gain on AA-Briefcase and a similar decline on GDPval-AA v2, plus 2-3 point drops on banking tool use, scientific Python, and long-context reasoning, per Artificial Analysis's benchmark report.

Long agent runs amplify these differences because one step’s output becomes the next step’s context, as rohanpaul_ai's analysis notes. The published results put that compounding on both sides of the ledger: fewer tokens can lower a task bill, but a higher token rate can still dominate it.

FrontierCode, WANDR, and migration

  • Merge-quality code: Cognition put Astra within 0.4 points of Fable 5 and 64% cheaper on FrontierCode 1.1. The benchmark uses 150 tasks with maintainers from 36 open-source repositories and zeros a run that misses a blocking requirement, as rohanpaul_ai's methodology summary explains.
  • Research and enterprise work: Perplexity reported Astra at 0.682 on WANDR, 13.5% above Fable 5.1 at 6.1% lower task cost. In an early preview, an early Box test reported 77% overall versus Sol’s 74% on its hardest enterprise set.
  • Migrations and reconstruction: ValsAI reported 68% on code migration, 10 points above second place, with 2-4x lower latency than comparable models. Epoch AI's prerelease evaluation put Astra at 169 ECI, above the prior 163 record but within Epoch’s uncertainty band, and its 46.7% MirrorCode score between Opus 4.7 and Fable 5. The same Epoch AI thread says Astra solved 2 of 68 Lean-formalized Erdős problems, where no earlier model had solved one.

Daybreak access

The launch page and full benchmark table briefly appeared, then disappeared, as synthwavedd's post recorded. The published OpenAI release now lists gpt-6-astra for the API and AWS, with access expanding from a limited group to ChatGPT Plus, Pro, Business, and Enterprise.

The production model refuses advanced cyber tasks such as proof-of-concept exploit creation; OpenAI says Daybreak will later extend less restrictive safeguards for defensive validation and malware analysis. Before release, sama's note said Astra had been done training for a while and that OpenAI was pacing follow-on models around safety and alignment work.

Recurrent depth

Separate reporting has tied Astra’s token efficiency to “recurrent depth,” a claim steph_palazzolo's report said needed more nuance in a subsequent follow-up. OpenAI’s launch post names pre-training, reinforcement learning, and alignment research, but does not specify a looped-transformer architecture.

In the public technical discussion, rasbt's explainer describes a looped transformer as reusing a layer stack for another pass over a hidden state. That adds internal compute without adding a second copy of the weights, and by itself does not remove visible chain-of-thought tokens.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR4 posts
$4.72 coding task2 posts
$10/$50 trade3 posts
FrontierCode, WANDR, and migration5 posts
Daybreak access2 posts
Recurrent depth2 posts
Share on X