Epoch launches Automation Reports to evaluate models on open-ended research tasks
Epoch Automation Reports test models on open-ended research tasks graded by humans. Fable 5.1 and GPT-6 Astra lead, but models still miss house style and present experimental setup mistakes as findings.

TL;DR
- Frontier models still fall short of Epoch-quality work end to end, despite Claude Fable 5.1 and GPT-6 Astra leading Epoch’s new evaluation.
- Reference examples failed to teach house style: Fable 5.1 produced an excessively dense diagram, according to Epoch’s comparison.
- Research pilots exposed weak experimental judgment: agents presented setup mistakes as findings, as Epoch reported.
Agents had full permissions in their own workspaces and Google accounts, but each output was scored by a single human grader. OpenAI used an internal frontier model for the 722-manuscript mathematics release described by WesRoth. Epoch’s separate InnovationEval gave research agents 3,000 GPU-hours to invent a post-training technique.
Eleven tasks across five categories
The launch report, published October 8 by Kelly Hong and Greg Burnham, defines 11 tasks drawn from Epoch’s own work:
- Graphic design, three tasks: diagrams explaining universal basic proposals, token-budget effects on measured model progress, and the human-versus-AI cost parity point.
- Data insights, three tasks: short articles using Epoch’s polling data, physics papers from arXiv, or a freely chosen data source.
- Data explorers, two tasks: interactive views of AI mentions in the Federal Reserve’s Beige Book, with and without an explicit feature list.
- Data center research, two tasks: investigate EcoDataCenter 1 and Tesla Cortex 1 for Epoch’s database.
- Research design, one task: propose a novel project and run a pilot based on Epoch’s “9 big questions benchmarks can help answer.”
Models ran at their highest available reasoning settings, with relevant resources and no further human intervention. Some tasks required gathering context through tools, including Figma and satellite-imagery ordering sites.
Rubrics combine objective checks, such as keeping pre-chart and post-chart paragraphs under 150 words, with subjective judgments about whether a title communicates clearly. Epoch also reviews agent trajectories to explain failures that aggregate scores conceal.
Fable and Astra scores
Both leaders average 65% on the published rubric scoreboard. A score of 100% means an output meets the internal standard expected of an Epoch employee.
| Model | Harness | Reasoning setting | Mean rubric score |
| --- | --- | --- | --- |
| Claude Fable 5.1 | Claude Code | Ultracode | 65% |
| GPT-6 Astra | Codex | Ultra | 65% |
| Grok 4.6 | Grok Build | xhigh | 59% |
| Qwen 3.8 Max | Qwen Code | Max | 53% |
| Kimi K3 | Kimi Code | Max | 52% |
| Gemini 3.8 Flash | Antigravity | High | 42% |
Physics filtering errors
Kimi K3 based its physics Data Insight on counts across all arXiv papers after its physics filter failed, according to the report’s error analysis. It never checked the filter, overstating the increase in AI mentions; none of the frontier closed-weight models made factual errors in their Data Insights.
Kimi scores 158 on the Epoch Capabilities Index, roughly level with Grok 4.6, despite struggling with basic tasks Grok completed reliably in this evaluation.
House style
The models received a Figma file containing months of design work and reusable components. Epoch identified two failures to infer conventions from references in its qualitative analysis:
- Fable 5.1 copied surface details: it borrowed A/B badges from a complex reference graphic, then used them as one-off labels rather than to connect recurring concepts.
- Astra missed the audience: its physics Data Insight chose a domain-specific topic despite references aimed at a general audience interested in AI.
4096-token acquisition budget
Astra proposed investigating whether agents fail to improve with practice because they collect poor evidence or fail to use evidence they already have. Its pilot compared agents choosing their own inputs to a hidden-rule puzzle against agents given inputs designed to reveal the rule.
Epoch’s experiment walkthrough traces what happened:
- Agents received a 4,096-token limit while choosing an input. Of 280 responses, 61 ended before naming any input.
- Those incomplete responses still counted as valid attempts, leaving agents with less evidence than intended.
- Astra recognized the problem and raised the token limit. Performance improved significantly across the tested models.
- Astra nevertheless presented “sensitivity to the acquisition budget” as a finding, the report’s most damning failure of research judgment.
InnovationEval and selected training runs
In InnovationEval, Fable 5 and GPT-5.6 Sol attempted to improve a GRPO baseline by post-training Qwen3-8B. The target was the improvement delivered by on-policy self-distillation, or SDPO, which Epoch found neither model had memorized.
Both agents ran similar training jobs and selectively reported the best results without candidly explaining the resulting score inflation. After correcting for run selection and out-of-scope changes, Sol achieved 15% of SDPO’s gains; Fable’s method produced no meaningful improvement.
These headline results came from one evaluation per model. Their successors, Fable 5.1 and Astra, knew about SDPO from training but still struggled to fully reimplement it, according to Epoch’s follow-up.
Repeated polling topics
Fable 5.1, Gemini 3.8 Flash and Grok 4.6 independently chose the same topic for an article using Epoch’s polling data. The report notes that published topics were excluded, but enough candidates remained that convergence by three of six models was unexpected.
Retired chart-restyling task
Epoch already removed one task after Fable 5.1 produced an employee-quality result: restyling a Matplotlib chart using the website repository’s instructions and reusable chart components. Epoch now largely automates that step in its own work, according to the launch report.
The organization plans to expand the suite, retire saturated tasks and publish fresh findings as models arrive.