Astra's Elo falls in Peter Gostev's repeated-play agent chess test
Astra's Elo fell during Peter Gostev's chess test, where agents can take notes and choose opponents. Later results suggest Opus improved, and GPT-6.1 Sol and Fable were added to the benchmark.

TL;DR
- Astra's Elo fell across 200 games in Peter Gostev's initial results, despite freedom to choose difficulty and take notes.
- Opus' early gain looked potentially random; Gostev later reported signs of improvement for Opus 5.5.
- The test runs agents in their native harnesses and expanded to GPT-6.1 Sol and Fable 5.1 in a later update.
Stockfish can deliberately pick suboptimal moves to provide weaker opponents. Gostev's live dashboard records each game's opponent, colour, ending, move count and post-game Elo.
The 200-game goal
Gostev gives each agent a /goal of learning and improving at chess, with as little operator input as possible. The rules leave several decisions to the agent:
- Play 200 games against Stockfish.
- Choose the opponent's difficulty.
- Take notes and devise a learning strategy.
- Avoid chess-engine assistance, which counts as cheating.
Runs use the agents' native harnesses, which Gostev identifies as Codex and Claude Code.
Astra's completed run
Gostev reported declining Elo for Astra in his initial post and called Opus' small gain potentially random. The two runs started at different times, with Astra beginning sooner.
Agent-written learning strategies
An unforgiving test for self-scaffolding: Gostev wants to see whether agents can simulate continuous learning through processes they devise themselves.
He described their freedom bluntly: “It can do what it wants,” he said in a reply. When commenters proposed teaching the agents a strategy, he declined to supply one.
Opus rises, Sol and Fable join
Opus 5.5 subsequently appeared to find a way to increase Elo, according to Gostev. His update added two more models:
- GPT-6.1 Sol.
- Claude Fable 5.1.
Gostev also re-deployed the dashboard after the original link broke.
Stockfish strength controls
Stockfish exposes separate skill-level and target-Elo controls in its official UCI documentation:
Skill Level: ranges from 0 to 20, with 20 as the default.UCI_LimitStrength: enables weaker play targeting the rating set throughUCI_Elo; it overridesSkill Level.UCI_Elo: its rating calibration uses a time control of 120 seconds plus a one-second increment, anchored to CCRL 40/4.
Because agents choose their opponent level, their runs can follow different difficulty schedules.