Peter Gostev reports Opus gains roughly 250 Elo in agent chess test
Peter Gostev reports Opus improving from roughly 1500 to 1750 Elo in a test where agents choose Stockfish difficulty and retain notes. He acknowledges measurement error and says he checked for prohibited cheating.

TL;DR
- Opus gained roughly 250 Elo: benchmark creator Peter Gostev reported a rise from about 1,500 to 1,750 during an agent chess test.
- Agents get 200 games and choose their own Stockfish difficulty and note-taking approach under the original rules.
- The starting rating is noisy: Gostev acknowledged measurement error when discussing the chart’s early spikes.
The live dashboard logs illegal moves alongside game results. A later screenshot includes a separate long-context Sol run that had completed just four games.
Opus gains roughly 250 Elo
Gostev initially described Opus’s improvement as tiny and potentially random in his first report. His later estimate put the gain at roughly 1,500 → 1,750 Elo, +250 Elo.
Astra had already finished its 200 games with a negative trend, while Opus was still playing. The apparent improvement changed substantially as the run progressed.
200 games in native harnesses
Gostev gives each agent a /goal to learn and improve at chess. The assignment leaves the learning loop largely to the agent:
- Play 200 games against Stockfish.
- Choose the opponent’s difficulty.
- Take notes and develop its own scaffolding.
- Avoid using a chess engine to generate its moves.
Codex and Claude Code provide the native harnesses, with as little input from Gostev as possible. He leaves the approach to the agent, arguing that teaching it would change the test.
The experiment evaluates the model-harness pair’s ability to build a learning workflow, alongside its chess play.
Opponent choice and rating noise
Stockfish exposes two documented strength controls:
Skill Level: a setting from 0 to 20.UCI_Elo: a target strength enabled byUCI_LimitStrength, which overridesSkill Level.
The target Elo is calibrated at a 120s+1s time control and anchored to CCRL 40/4. Elo also depends on match conditions and the opponent pool, according to Stockfish’s FAQ.
Gostev acknowledged measurement error and suggested excluding the first few spikes when estimating starting strength. Those early chart peaks are a shaky baseline for calculating improvement.
Context windows and new entrants
Growing notes and files might strain Astra’s context window, Gostev suggested as one possible explanation for its deterioration. He proposed trying a larger window.
Sol subsequently appeared in two configurations in the later leaderboard:
- Standard Sol: 258K context, 19 completed games.
- Sol, long context: 828K context, four completed games.
Opus 5.5 and Fable 5.1 displayed 1M context, and every listed run used max reasoning. The longer-context Sol run was still too early to establish a comparable 200-game trajectory.
Engine checks and illegal moves
Gostev said he had checked Opus 5.5 and found it playing fairly. He would invalidate a score if prohibited engine use were discovered.
The dashboard separately records illegal-move counts and percentages, a rule-following measure alongside the Elo estimate.