OpenAI's GeneBench-Pro Exposes That Top AI Fails Real Biology 70% of the Time
OpenAI's new GeneBench-Pro tests whether AI agents can handle the messy judgment calls of real computational biology research — and current models solve fewer than a third of problems.

- GeneBench-Pro launches: OpenAI releases a 129-problem research-level benchmark testing AI agents on messy, judgment-heavy computational biology tasks.
- Best score is 31.5%: GPT-5.6 Sol (Pro mode) leads all models; GPT-5 scored below 5% when the original GeneBench was first built.
- Open-source release: 10 representative problems with data files, ground-truth answers, and a reference grader are available on Hugging Face under CC-BY-4.0.
- Key failure mode: Models spot data issues but don't act on them — they notice diagnostic signals and then pick the wrong analysis path anyway.
- Massive cost gap: Human experts need 20–40 hours per problem (~$4,000–$8,000); AI inference costs only a few dollars per problem.
- Independent evaluation coming: A 50-question subset will go to Artificial Analysis for third-party benchmarking; full benchmark may be saturated by year-end.
Most AI benchmarks in biology test whether a model knows facts or can run a standard pipeline. GeneBench-Pro tests something harder: can an agent look at a messy, real-world genomics dataset, figure out what is wrong with it, choose the right analysis strategy, and arrive at a conclusion that a researcher could actually act on? That is a very different bar, and today's best models clear it less than a third of the time.
The benchmark gap nobody was measuring
Existing biology benchmarks mostly measure knowledge retrieval, execution of routine pipelines, or a single analysis step. They do not capture what actually occupies most of a computational scientist's time: cleaning and normalizing data, exploratory analysis, statistical model selection, diagnostic iteration, and producing a conclusion that informs a downstream scientific or translational decision.
GeneBench-Pro is a challenging, research-level benchmark for testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires. It is the successor to the original GeneBench, and it deliberately raises the difficulty floor. OpenAI calls the target skill "research taste" , the chains of judgment calls that shape an analysis: which questions the data can actually support, how early warning signs should change your model, and when your initial plan needs to be thrown out.
129 problems, no shortcuts allowed
GeneBench-Pro covers 129 questions across genomics, quantitative biology, and translational medicine, capturing the complexity, iterative nature, and ambiguity of scientific research in computational biology. The 10 domains include:
- Statistical and population genetics
- Regulatory omics and functional genomics
- Proteomics and biomarkers
- Clinical variant interpretation and pharmacogenomics
- Cancer somatic genomics and liquid biopsy
- Microbial and forensic genetics
What makes the benchmark unusually rigorous is how each problem is constructed. Each GeneBench-Pro problem is built synthetically: the full causal structure is known and the data-generating process is directly simulated. That enables tuning the complexity of each problem, ensuring that reasonable differences in subjective analytical choices still produce accepted numerical results, and verifying through ablation studies that plausible but incorrect analyses fail. In other words, you cannot game it by picking a defensible-but-wrong path.
Each agent gets access to an isolated workspace with a short prompt, data files, and a standard bioinformatics stack including Python, scientific computing libraries, and tools like PLINK 2.0. The answer must be returned as a structured JSON object, and correctness is graded deterministically against known targets , no rubric-based scoring, no verbosity bonuses.
What the numbers actually say
The strongest model, GPT-5.6 Sol, attains a pass rate of 28.7% at the highest reasoning level, rising to 31.5% with Pro mode enabled. That is a sharp increase from when OpenAI began building the original GeneBench; at that time, the best frontier model, GPT-5, scored below 5%.
The scaling story is also striking. At the lowest reasoning level, GPT-5.6 Sol only achieves a single-digit pass rate. At the highest reasoning level, GPT-5.6 Sol solves nearly six times as many questions as GPT-5.2 does while using about two-thirds as many tokens. More thinking budget, used efficiently, matters enormously here.
The strongest external baseline, Gemini 3.1 Pro, achieves 11.2%. Comparisons across model families suggest that the performance gap between GPT-5.6, GPT-5.5, and leading open-source models is significantly larger than would be expected when extrapolating from coding benchmarks, indicating that open-source models are more specialized for coding than for broader reasoning ability.
Where models still fall apart
Models often complete substantial portions of the workflow but exhibit a consistent gap between noticing and acting: they identify local diagnostic signals but fail to propagate the implication to the corresponding analysis decision, and as a result select wrong estimators or persist on initially plausible but incorrect analysis paths.
One expert reviewer put it plainly: "Most of the agents failed on data discrepancies. They aren't cautious enough about data issues. And a lot of biological data has irregularities." The contrast between GPT-5.5 and GPT-5.6 Sol on a pharmacogenomics problem illustrates this well , GPT-5.5 fits a conventional Cox survival model but misses treatment-confounder feedback entirely, while GPT-5.6 Sol correctly applies a marginal structural model with inverse-probability weighting to handle the same issue.
The economic argument
In a survey, reviewers estimated that a typical GeneBench-Pro problem would take a human expert around 20 to 40 hours to complete. At a conservative $200 per hour, that puts the human labor cost of a single problem in the thousands of dollars. Current AI agents are still too unreliable to replace human experts, but inference costs sit at only several dollars per problem. Even partial automation at current capability levels represents a significant cost reduction for biobank-scale research programs.
Sequencing costs have plummeted, and biobank-scale datasets now link molecular, phenotypic, and health-record information at unprecedented breadth. The limiting factor is shifting from data generation to turning information into actionable insights. GeneBench-Pro is designed to measure progress against exactly that bottleneck.
What is open and how to use it
OpenAI is fully open-sourcing 10 representative problems on Hugging Face under a CC-BY-4.0 license, with complete data files, PDF case studies, ground-truth answers, and a reference grader. A 50-question subset will also be provided to Artificial Analysis for independent third-party benchmarking. The public package is free to use. Running a problem is straightforward:
python3 reference_grader.py \
problems/multiparent_qtl_hmm_lmm/eval_config.json \
path/to/eval_answer.jsonEach eval_config.json contains the task prompt, staged data files, answer schema, ground-truth values, and grader tolerances. The boolean passed field is the authoritative grading decision.
The benchmark is most useful if you are:
- Evaluating frontier models or fine-tunes on scientific reasoning, not just code execution
- Building agentic pipelines for genomics or translational research workflows
- Researching where and why LLM agents break down on multi-step quantitative analysis
- Tracking progress toward AI-assisted hypothesis triage and target prioritization in drug discovery
At the current pace, this benchmark may be saturated by the end of the year. That is either an optimistic sign of how fast the field is moving, or a reminder that the community needs to keep building harder evals as fast as models improve. Either way, GeneBench-Pro sets a new standard for what "can it do biology?" actually means.