Skip to content
AI Primer
workflow

Astra coordinating Opus 5.5 wins one coding test on quality

Kevin Kern's single-task comparison favored Astra coordinating Opus 5.5 for quality. Solo runs were close, cost less and finished about three times faster; a rerun corrected for polling costs.

4 min read
Astra coordinating Opus 5.5 wins one coding test on quality
Astra coordinating Opus 5.5 wins one coding test on quality

TL;DR

  • An Astra-led run with Opus 5.5 building and Astra reviewing topped Kevin Kern’s single-task test; Kern’s rerun used an MCP completion wait, and his updated results score it 83.3/100 at an estimated $31.10.
  • Solo Opus remained close at less than half that estimated cost, as Kern’s original comparison suggested; the updated table puts it at 79.8/100 and $14.93.
  • Status checks consumed $9.75 in an earlier coordinated run, according to Kern’s corrected cost chart. The completion-wait rerun recorded no identified polling cost.
  • Solo runs finished substantially sooner; Kern’s timing post described them as roughly three times faster in his initial comparison.

The run notes include a polished design from a costly Luna-supervised setup, plus a review that missed a false “saved” response. Kern also couldn’t make Opus the lead agent in Codex Desktop using his Claude subscription, despite wanting to keep its app-thread spawning and browser tools.

One TypeScript feature, several role splits

Kern gave the same feature and demo data from his finance app to solo agents, split implementations, and build-then-review pairs. His published methodology and run notes weight code at 50%, UX at 25%, and design at 25%; this is one private task, not a general model ranking.

The runs used native Codex and Claude Code rather than a custom harness, Kern said in a reply. Some pairs split backend and UI work; others had one model build everything before the second reviewed it, as he clarified in another reply.

The updated leaderboard

Kern’s live results table lists estimated API token costs at list prices:

  1. Astra → Opus → Astra review, MCP wait: 83.3/100, $31.10, 1h 17m.
  2. Opus → Astra review: 81.5/100, $32.48, 1h 24m.
  3. Opus solo: 79.8/100, $14.93, 43m 25s.
  4. Astra solo: 76.5/100, $14.48, 35m 55s.

The lead over solo Opus was 3.5 points for just over twice the estimated spend. The attached charts in Kern’s rerun post and his initial comparison show earlier score snapshots; the figures above follow the updated page.

A $9.75 polling bill

In the earlier Astra → Opus → Astra review run, identified status polling accounted for $9.75 of its $43.22 estimated cost, or 22.5%, in Kern’s revised breakdown. The later MCP run let the parent wait for Opus to complete, then resume review. It scored 83.3 at $31.10 with no identified polling cost.

The $12.12 difference cannot all be assigned to the wait: implementation and review differed between runs, Kern says in his run notes. A reply about native multiagent tools explains the mechanism: parking the parent outside the code-mode polling loop avoids repeated model-driven checks.

Luna’s design score and the runtime gap

Luna supervising an Astra → Opus build produced the best design score, 86/100, but just 66/100 for code quality at an estimated $70.41, according to the component rankings. Kern called the middle layer unnecessarily expensive. Solo Luna cost $0.89 and scored 66 overall, but its recorded elapsed time was 13h 45m, an outlier to Kern’s faster-solo observation.

A separate Jev router trial

Kern also tested a smaller coding task with a router switching between Codex models. Astra solo completed 14/14 checks in 21.3 minutes for $6.98; the Jev-routed team reached its 30-minute limit with 12/14 checks at $6.72. These are Kern’s separate trial, not additional runs of the finance-app feature.

Codex Desktop’s lead-agent constraint

Kern’s preferred arrangement was Opus leading with Astra reviewing, but his Codex Desktop account says he couldn’t select Opus as the main agent there using his Claude subscription. Moving to Claude Code would lose the app thread spawning, native browser use, and computer use he relied on; his thread-list screenshot shows the parallel app work behind that constraint.

A passing run with a false “saved” response

The feature-check table and code samples show that the winning MCP run passed 15/15 checks, yet its review missed a retry bug. When a project-creation request reused an existing ID with a different name and budget, the handler returned status: "saved" alongside the old project, without applying the submitted values.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 7 threads
TL;DR2 posts
One TypeScript feature, several role splits3 posts
The updated leaderboard1 post
A $9.75 polling bill2 posts
Luna’s design score and the runtime gap1 post
A separate Jev router trial1 post
Codex Desktop’s lead-agent constraint2 posts
Share on X