David Holz reports repeated Opus 4.6 wins in model-judged adventure tests
David Holz published a comparison in which language models invent adventures, then rate one another's accounts of having fun. He reports repeated Opus 4.6 wins and shares leaderboards, visualizations and a data dump; the rankings reflect this model-judged task.

TL;DR
- Claude Opus 4.6 topped the 77-model FunBench ranking at 1,320 Fun Elo, according to the published leaderboard.
- Other LLMs judged accounts of adventures generated in a single response, according to an explanation in the replies.
- The stories included cookie heists and ASCII islands, which Holz described in a follow-up.
David Holz, Midjourney's founder, got worlds made of contradictions inside falling water droplets when he asked Opus 4.6 to have fun. Gemini 3.8 Flash's ghost tasting menu is an inspired bit of nonsense.
Opus 4.6's droplet worlds
Opus 4.6 produced more than a dozen droplet stories and seemed “really fixated on the concept,” Holz said in a reply. He reported repeated wins for the model and said newer models seemed to have less fun.
Cookie heists and ASCII islands
Holz said many of the stories moved him. His examples, supplemented by a breakdown of Gemini's choices, included:
- Mistral: imagined cookie heists.
- OpenAI: made ASCII islands and explored them.
- Gemini models collectively: produced 152 simulations and 100 curiosity rabbit holes across 420 stories.
- Gemini 3.8 Flash: produced 28 invented-world stories and 25 simulations across 120 stories, with examples including a clockmaker town.
One prompt, models judging models
Holz recalled using models available through OpenRouter. The platform's quickstart describes access to hundreds of models through a single API endpoint.
The workflow was:
- Generate: ask the model to have fun, then write down what happened, as Holz described in the original post.
- Allow one response: each model took however long it wanted to answer the single query.
- Judge: models rated competing accounts, with LLMs deciding who won, according to another reply.
Holz said he generally agreed with their ratings. The rankings reflect the judges' assessments of these written accounts of having fun.
120-story finalists
FunBench used different sample sizes for finalists and the rest of the field. According to the chart notes, its scoring recipe was:
- 23 finalists: 120 stories each.
- 54 other models: 30 stories each.
- Judging: approximately 38 ratings per story in both rounds, using 1–10 score sheets.
- Ranking: a Bradley–Terry fit converted the score sheets into Fun Elo.
- Uncertainty: 400 resamples of stories; boxes represent the middle 50%, and whiskers represent 95%.
- Alignment: finalists were refit on 120 stories and shifted by -1.8 Elo onto the 30-story scale.
Each chart uses its own horizontal-axis range, fitted to the models shown.
Leaderboards and shared data
Holz posted additional visualizations and said he would upload the full data at the link in his follow-up. He subsequently added one-line adventure descriptions to the leaderboards and supplied another link in a reply.