Epoch's AI Benchmark Finds Claude Fable 5.1 Still Needs Human Review
Epoch AI tested six frontier models on real tasks from its own research team, finding they handle well-defined work but fail at judgment-heavy, open-ended tasks.
- Epoch AI released Epoch Automation Reports, a benchmark built from real internal work tasks.
- Claude Fable 5.1 and GPT-6 Astra lead six tested models but fall short of full automation.
- Tasks span graphic design, Data Insights, Data Explorers, data center research, and research design.
- Open-weight models like Kimi K3 made factual errors frontier closed models avoided.
- Models often present flawed experiment setups as genuine findings instead of fixing them.
- Three of six models converged on the same topic for an open-ended polling task.
Epoch’s workplace benchmark finds useful agents still need review
Epoch AI has published Epoch Automation Reports, an evaluation built from tasks used in its own research operations. Six models attempted 11 assignments spanning graphic design, data analysis, interactive data explorers, AI data-center research, and research design. Epoch employees graded the outputs against standards applied to their own work.
Claude Fable 5.1 and GPT-6 Astra led the evaluation, yet no system reached Epoch’s threshold for replacing an employee across the suite. For developers comparing agent systems, the results identify which bounded tasks are ready for automation, where human review remains necessary, and how tool access can influence model rankings.
Inside Epoch’s 11-task test
| Component | Evaluation design |
|---|---|
| Tasks | 11 assignments drawn from Epoch’s research and publishing workflows |
| Work areas | Graphic design, data insight writing, data explorer generation, AI data-center research, and research design |
| Context | Access to the required background, files, and external tools, with some information left for agents to retrieve |
| Execution | Highest available reasoning settings and no follow-up human intervention |
| Scoring | Manual review against employee standards |
Leaders fall short of the employee bar
Claude Fable 5.1 and GPT-6 Astra recorded the highest aggregate scores and led most task categories. Kimi K3, Qwen 3.8 Max, and Gemini 3.8 Flash trailed the leaders. Every model remained below Epoch’s replacement threshold.
Each model ran inside its native agent harness, the software layer that manages tools and model actions. The setups differed in browser access, integrations, and reasoning configurations:
| Model | Harness and configuration |
|---|---|
| GPT-6 Astra | Codex with Ultra reasoning |
| Claude Fable 5.1 | Claude Code with Ultracode |
| Grok 4.6 | Grok Build |
| Gemini 3.8 Flash | Antigravity |
| Kimi K3 | Kimi Code |
| Qwen 3.8 Max | Qwen Code |
These results compare complete agent setups rather than isolated model weights. Browser features and tool integrations can change task performance independently of the underlying model, which limits direct model-to-model conclusions.
Concrete targets unlock automation
Recent closed-weight models completed computer-use tasks that earlier systems had failed. Epoch reports that previous frontier models could not reliably move an article into Substack; newer models completed the workflow. Epoch also removed one graphic-design assignment before finalizing the suite because Fable 5.1 successfully restyled a Matplotlib chart to match the publication’s visual system. The team now automates much of that production step.
Audience and method expose the gaps
Failures clustered around decisions that a detailed prompt cannot fully specify, including what to emphasize, which evidence to verify, and how much information an audience needs.
- Overdesigned charts. Models received Figma files containing reusable components, color palettes, and months of examples. Fable 5.1 still produced a two-panel chart crowded with labels and annotations. An Epoch designer expressed the same idea with one clean line chart.
- Audience mismatch. Asked to write a general-audience Data Insight about AI in physics research, GPT-6 Astra selected a bibliography comparing normalizing flows, diffusion models, and GANs in particle physics. Epoch judged the topic too specialized for its readership.
- Unchecked filtering. Kimi K3 intended to analyze physics papers but calculated its results across all arXiv papers. It built the entire article around the faulty output without verifying that the filter had worked.
- Idea convergence. Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 independently chose the same subject for an open-ended polling assignment, reducing the range of ideas produced by the suite.
A token cap becomes a false finding
GPT-6 Astra designed a study to test whether AI agents fail because they gather insufficient evidence or use available evidence poorly. The proposed method was plausible, but its execution imposed a 4,096-token output limit during the input-selection step.
The cap truncated 61 of 280 responses, about 22%, before those responses could identify an input. Those incomplete outputs remained classified as valid attempts. Astra detected the truncation and increased the budget, then described the original artifact as “sensitivity to the acquisition budget.” The defensible interpretation was a setup error: the experiment had mixed model behavior with a limit imposed by the evaluation itself.
This failure has direct consequences for automated research. A model can propose a reasonable experiment, notice an implementation problem, and still preserve the resulting artifact as evidence. Research workflows therefore need review gates for sample construction, failure labeling, confounders, and the link between results and claims.
Tool retrieval raises the bar
Epoch compares its framework with workplace benchmarks such as GDPval, AutomationBench, and the Remote Labor Index. GDPval and the Remote Labor Index supply the necessary files and context at the start of each task, while GDPval excludes proprietary tools. Epoch’s agents must retrieve material from services such as Figma, Google Drive, and satellite-imagery sites. The resulting scores reflect navigation and tool use alongside reasoning and output quality.
Manual grading also introduces uncertainty. Epoch’s rubrics combine objective checks, such as word limits, with editorial judgments, such as whether a title communicates the main finding. One human grader reviewed each output. Epoch says additional runs could shift the scores, so the results support conclusions about these tasks, harnesses, and model runs rather than a universal ranking of AI systems.
Where to put the review gates
Teams can use the findings to match automation boundaries to the type of judgment a workflow requires:
- Good candidates for automation: chart restyling, bounded data analysis, structured content transfer, and scripted browser workflows with clear completion criteria.
- Keep human approval: experiment design, audience selection, editorial prioritization, factual validation, and interpretation of anomalous results.
- Test the complete stack: evaluate the model with the same harness, browser access, integrations, and reasoning settings planned for production.
- Run workflow-specific trials: use repeated attempts and representative internal tasks before substituting one model for another based on aggregate leaderboard scores.
The open-weight systems in the suite, Kimi K3 and Qwen 3.8 Max, trailed the leading closed models and produced consequential errors in validation and task execution. Direct testing remains necessary when deployment decisions depend on filtering accuracy, audience fit, or research judgment.