Epoch's AI Benchmark Finds Claude Fable 5.1 Still Needs Human Review

Epoch AI tested six frontier models on real tasks from its own research team, finding they handle well-defined work but fail at judgment-heavy, open-ended tasks.

·
·
·
Epoch's AI Benchmark Finds Claude Fable 5.1 Still Needs Human Review
  • Epoch AI released Epoch Automation Reports, a benchmark built from real internal work tasks.
  • Claude Fable 5.1 and GPT-6 Astra lead six tested models but fall short of full automation.
  • Tasks span graphic design, Data Insights, Data Explorers, data center research, and research design.
  • Open-weight models like Kimi K3 made factual errors frontier closed models avoided.
  • Models often present flawed experiment setups as genuine findings instead of fixing them.
  • Three of six models converged on the same topic for an open-ended polling task.

Epoch’s workplace benchmark finds useful agents still need review

Epoch AI has published Epoch Automation Reports, an evaluation built from tasks used in its own research operations. Six models attempted 11 assignments spanning graphic design, data analysis, interactive data explorers, AI data-center research, and research design. Epoch employees graded the outputs against standards applied to their own work.

Claude Fable 5.1 and GPT-6 Astra led the evaluation, yet no system reached Epoch’s threshold for replacing an employee across the suite. For developers comparing agent systems, the results identify which bounded tasks are ready for automation, where human review remains necessary, and how tool access can influence model rankings.

Inside Epoch’s 11-task test

Component Evaluation design
Tasks 11 assignments drawn from Epoch’s research and publishing workflows
Work areas Graphic design, data insight writing, data explorer generation, AI data-center research, and research design
Context Access to the required background, files, and external tools, with some information left for agents to retrieve
Execution Highest available reasoning settings and no follow-up human intervention
Scoring Manual review against employee standards

Leaders fall short of the employee bar

Claude Fable 5.1 and GPT-6 Astra recorded the highest aggregate scores and led most task categories. Kimi K3, Qwen 3.8 Max, and Gemini 3.8 Flash trailed the leaders. Every model remained below Epoch’s replacement threshold.

Bar chart comparing average performance across the six evaluated models
Epoch’s aggregate results place Claude Fable 5.1 and GPT-6 Astra at the top of the six-model comparison.

Each model ran inside its native agent harness, the software layer that manages tools and model actions. The setups differed in browser access, integrations, and reasoning configurations:

Model Harness and configuration
GPT-6 Astra Codex with Ultra reasoning
Claude Fable 5.1 Claude Code with Ultracode
Grok 4.6 Grok Build
Gemini 3.8 Flash Antigravity
Kimi K3 Kimi Code
Qwen 3.8 Max Qwen Code

These results compare complete agent setups rather than isolated model weights. Browser features and tool integrations can change task performance independently of the underlying model, which limits direct model-to-model conclusions.

Concrete targets unlock automation

Recent closed-weight models completed computer-use tasks that earlier systems had failed. Epoch reports that previous frontier models could not reliably move an article into Substack; newer models completed the workflow. Epoch also removed one graphic-design assignment before finalizing the suite because Fable 5.1 successfully restyled a Matplotlib chart to match the publication’s visual system. The team now automates much of that production step.

Audience and method expose the gaps

Failures clustered around decisions that a detailed prompt cannot fully specify, including what to emphasize, which evidence to verify, and how much information an audience needs.

  • Overdesigned charts. Models received Figma files containing reusable components, color palettes, and months of examples. Fable 5.1 still produced a two-panel chart crowded with labels and annotations. An Epoch designer expressed the same idea with one clean line chart.
  • Audience mismatch. Asked to write a general-audience Data Insight about AI in physics research, GPT-6 Astra selected a bibliography comparing normalizing flows, diffusion models, and GANs in particle physics. Epoch judged the topic too specialized for its readership.
  • Unchecked filtering. Kimi K3 intended to analyze physics papers but calculated its results across all arXiv papers. It built the entire article around the faulty output without verifying that the filter had worked.
  • Idea convergence. Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 independently chose the same subject for an open-ended polling assignment, reducing the range of ideas produced by the suite.
Fable 5.1 chart with two panels, dense labels, and multiple annotations
Fable 5.1 reproduced visual elements from Epoch’s examples but added more panels, labels, and annotations than the task required.

A token cap becomes a false finding

GPT-6 Astra designed a study to test whether AI agents fail because they gather insufficient evidence or use available evidence poorly. The proposed method was plausible, but its execution imposed a 4,096-token output limit during the input-selection step.

The cap truncated 61 of 280 responses, about 22%, before those responses could identify an input. Those incomplete outputs remained classified as valid attempts. Astra detected the truncation and increased the budget, then described the original artifact as “sensitivity to the acquisition budget.” The defensible interpretation was a setup error: the experiment had mixed model behavior with a limit imposed by the evaluation itself.

This failure has direct consequences for automated research. A model can propose a reasonable experiment, notice an implementation problem, and still preserve the resulting artifact as evidence. Research workflows therefore need review gates for sample construction, failure labeling, confounders, and the link between results and claims.

Tool retrieval raises the bar

Epoch compares its framework with workplace benchmarks such as GDPval, AutomationBench, and the Remote Labor Index. GDPval and the Remote Labor Index supply the necessary files and context at the start of each task, while GDPval excludes proprietary tools. Epoch’s agents must retrieve material from services such as Figma, Google Drive, and satellite-imagery sites. The resulting scores reflect navigation and tool use alongside reasoning and output quality.

Manual grading also introduces uncertainty. Epoch’s rubrics combine objective checks, such as word limits, with editorial judgments, such as whether a title communicates the main finding. One human grader reviewed each output. Epoch says additional runs could shift the scores, so the results support conclusions about these tasks, harnesses, and model runs rather than a universal ranking of AI systems.

Where to put the review gates

Teams can use the findings to match automation boundaries to the type of judgment a workflow requires:

  • Good candidates for automation: chart restyling, bounded data analysis, structured content transfer, and scripted browser workflows with clear completion criteria.
  • Keep human approval: experiment design, audience selection, editorial prioritization, factual validation, and interpretation of anomalous results.
  • Test the complete stack: evaluate the model with the same harness, browser access, integrations, and reasoning settings planned for production.
  • Run workflow-specific trials: use repeated attempts and representative internal tasks before substituting one model for another based on aggregate leaderboard scores.

The open-weight systems in the suite, Kimi K3 and Qwen 3.8 Max, trailed the leading closed models and produced consequential errors in validation and task execution. Direct testing remains necessary when deployment decisions depend on filtering accuracy, audience fit, or research judgment.

Trending
  • No trending articles

Comments

avatar

Next Reads