Skip to content
AI Primer
breaking

Astra's Elo falls in Peter Gostev's repeated-play agent chess test

Astra's Elo fell during Peter Gostev's chess test, where agents can take notes and choose opponents. Later results suggest Opus improved, and GPT-6.1 Sol and Fable were added to the benchmark.

2 min read
Astra's Elo falls in Peter Gostev's repeated-play agent chess test
Astra's Elo falls in Peter Gostev's repeated-play agent chess test

TL;DR

Stockfish can deliberately pick suboptimal moves to provide weaker opponents. Gostev's live dashboard records each game's opponent, colour, ending, move count and post-game Elo.

The 200-game goal

Gostev gives each agent a /goal of learning and improving at chess, with as little operator input as possible. The rules leave several decisions to the agent:

  • Play 200 games against Stockfish.
  • Choose the opponent's difficulty.
  • Take notes and devise a learning strategy.
  • Avoid chess-engine assistance, which counts as cheating.

Runs use the agents' native harnesses, which Gostev identifies as Codex and Claude Code.

Astra's completed run

Gostev reported declining Elo for Astra in his initial post and called Opus' small gain potentially random. The two runs started at different times, with Astra beginning sooner.

Agent-written learning strategies

An unforgiving test for self-scaffolding: Gostev wants to see whether agents can simulate continuous learning through processes they devise themselves.

He described their freedom bluntly: “It can do what it wants,” he said in a reply. When commenters proposed teaching the agents a strategy, he declined to supply one.

Opus rises, Sol and Fable join

Opus 5.5 subsequently appeared to find a way to increase Elo, according to Gostev. His update added two more models:

  • GPT-6.1 Sol.
  • Claude Fable 5.1.

Gostev also re-deployed the dashboard after the original link broke.

Stockfish strength controls

Stockfish exposes separate skill-level and target-Elo controls in its official UCI documentation:

  • Skill Level: ranges from 0 to 20, with 20 as the default.
  • UCI_LimitStrength: enables weaker play targeting the rating set through UCI_Elo; it overrides Skill Level.
  • UCI_Elo: its rating calibration uses a time control of 120 seconds plus a one-second increment, anchored to CCRL 40/4.

Because agents choose their opponent level, their runs can follow different difficulty schedules.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR3 posts
The 200-game goal2 posts
Astra's completed run1 post
Agent-written learning strategies2 posts
Opus rises, Sol and Fable join1 post
Share on X