Skip to content
AI Primer
update

Peter Gostev reports Opus gains roughly 250 Elo in agent chess test

Peter Gostev reports Opus improving from roughly 1500 to 1750 Elo in a test where agents choose Stockfish difficulty and retain notes. He acknowledges measurement error and says he checked for prohibited cheating.

3 min read
Peter Gostev reports Opus gains roughly 250 Elo in agent chess test
Peter Gostev reports Opus gains roughly 250 Elo in agent chess test

TL;DR

  • Opus gained roughly 250 Elo: benchmark creator Peter Gostev reported a rise from about 1,500 to 1,750 during an agent chess test.
  • Agents get 200 games and choose their own Stockfish difficulty and note-taking approach under the original rules.
  • The starting rating is noisy: Gostev acknowledged measurement error when discussing the chart’s early spikes.

The live dashboard logs illegal moves alongside game results. A later screenshot includes a separate long-context Sol run that had completed just four games.

Opus gains roughly 250 Elo

Gostev initially described Opus’s improvement as tiny and potentially random in his first report. His later estimate put the gain at roughly 1,500 → 1,750 Elo, +250 Elo.

Astra had already finished its 200 games with a negative trend, while Opus was still playing. The apparent improvement changed substantially as the run progressed.

200 games in native harnesses

Gostev gives each agent a /goal to learn and improve at chess. The assignment leaves the learning loop largely to the agent:

  • Play 200 games against Stockfish.
  • Choose the opponent’s difficulty.
  • Take notes and develop its own scaffolding.
  • Avoid using a chess engine to generate its moves.

Codex and Claude Code provide the native harnesses, with as little input from Gostev as possible. He leaves the approach to the agent, arguing that teaching it would change the test.

The experiment evaluates the model-harness pair’s ability to build a learning workflow, alongside its chess play.

Opponent choice and rating noise

Stockfish exposes two documented strength controls:

  • Skill Level: a setting from 0 to 20.
  • UCI_Elo: a target strength enabled by UCI_LimitStrength, which overrides Skill Level.

The target Elo is calibrated at a 120s+1s time control and anchored to CCRL 40/4. Elo also depends on match conditions and the opponent pool, according to Stockfish’s FAQ.

Gostev acknowledged measurement error and suggested excluding the first few spikes when estimating starting strength. Those early chart peaks are a shaky baseline for calculating improvement.

Context windows and new entrants

Growing notes and files might strain Astra’s context window, Gostev suggested as one possible explanation for its deterioration. He proposed trying a larger window.

Sol subsequently appeared in two configurations in the later leaderboard:

  • Standard Sol: 258K context, 19 completed games.
  • Sol, long context: 828K context, four completed games.

Opus 5.5 and Fable 5.1 displayed 1M context, and every listed run used max reasoning. The longer-context Sol run was still too early to establish a comparable 200-game trajectory.

Engine checks and illegal moves

Gostev said he had checked Opus 5.5 and found it playing fairly. He would invalidate a score if prohibited engine use were discovered.

The dashboard separately records illegal-move counts and percentages, a rule-following measure alongside the Elo estimate.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR4 posts
Opus gains roughly 250 Elo1 post
200 games in native harnesses3 posts
Opponent choice and rating noise1 post
Context windows and new entrants2 posts
Engine checks and illegal moves1 post
Share on X