OpenAI Launches GeneBench-Pro to Test AI Judgment on Complex Biological Data

OpenAI Launches GeneBench-Pro to Test AI Judgment on Complex Biological Data

A new 129-problem benchmark measures how AI agents handle messy genomics and computational biology tasks requiring multi-stage scientific reasoning.

Picture a computational biologist staring at a noisy genomics dataset riddled with sequencing errors, hidden confounders and ancestry misclassifications. The question isn’t simply “what does this data say?” — it’s “what *can* this data actually support, and how do I know when my analysis plan needs to change?” That kind of iterative, judgment-heavy scientific work has long been considered beyond the reach of AI. OpenAI is now trying to measure exactly how close its models are getting.

The company has announced GeneBench-Pro, a research-level benchmark designed to test whether AI agents can handle the kind of messy, multi-stage biological analysis that real computational scientists do every day. It’s harder, more realistic, and considerably more demanding than its predecessor, the original GeneBench.

What GeneBench-Pro Actually Tests

The benchmark comprises 129 problems spread across 10 primary domains and 21 subdomains, covering areas including statistical genetics, cancer genomics, pharmacogenomics and forensic genetics. Each problem hands the AI agent a brief experimental context, a noisy or imperfect dataset, and a specific quantity to estimate — what the researchers call the “target estimand” — that would matter for a real downstream decision.

That’s quite different from asking a model to recall a fact or solve a single equation. The agent must choose an analysis path, run diagnostics, recognise when the data can’t support the original question, and sometimes pivot entirely to an alternative approach. It’s designed to mirror how a postdoctoral researcher or senior scientist actually works through a problem, not how a student answers a textbook exercise.

One of the more technically interesting design choices is the use of synthetically generated data with fully known causal structures. Because the ground truth is specified in advance, grading is deterministic — there’s no reliance on a judge’s subjective rubric. The benchmark can also verify that plausible-but-wrong analysis paths genuinely fail, rather than accidentally scoring well.

OpenAI has open-sourced 10 representative problems for public inspection. The full 129-problem suite is used internally for systematic model evaluation.

The Numbers — and the Gap

The headline result is that OpenAI’s strongest model, GPT‑5.6 Sol, achieves a 28.7% pass rate at the highest reasoning level across the full benchmark, rising to 31.5% when “Pro mode” is enabled. Those figures are better than anything else currently on the leaderboard — but they also mean the model gets the answer wrong nearly seven times out of ten.

On top of that, the gap between models is striking. According to technology coverage and community reporting — figures not yet independently confirmed on OpenAI’s official evaluation pages — Claude Opus 4.8 from Anthropic scores around 16.0%, while Google DeepMind’s Gemini 3.1 Pro comes in at close to 3.1%. That’s a wide spread for what are all considered frontier-class models.

It’s a humbling set of numbers.

OpenAI suggests that at the current pace of improvement, GeneBench-Pro could be “saturated” — meaning top models approach near-perfect performance — around the end of 2026. Whether that timeline holds depends on how quickly the underlying models improve on multi-step scientific reasoning, which has historically been harder to accelerate than raw knowledge recall.

“Research Taste” as a Benchmark Target

OpenAI describes what GeneBench-Pro is trying to capture as “research taste” — a phrase that sounds informal but refers to something quite specific. It means the sequence of judgment calls a scientist makes: which questions the available data can support, how diagnostic results should alter the model or the estimand, and when there’s enough evidence to commit to a conclusion.

Prior AI benchmarks in biology have tended to use clean, curated datasets and relatively straightforward statistical tasks. Critics argued those tests didn’t reflect the ambiguity and iteration that real genomics research involves. GeneBench-Pro pushes back against that by deliberately introducing noise, biases, quality-control failures and hidden confounders — the kind of problems that make biological data genuinely difficult to work with.

External domain experts, including graduate students, postdoctoral researchers and professors, were consulted to validate that the problems are realistic and that the intended answers are actually recoverable from the provided data.

Not everyone is convinced that high benchmark scores translate cleanly into safe clinical deployment, mind. Some AI and bioethics researchers caution that synthetic-data benchmarks, however carefully designed, may miss complexities present in real patient data. And a benchmark created and administered by the same company whose model tops the leaderboard raises questions about independent scrutiny — a concern the AI research community has raised about proprietary evaluations before.

What This Means for Kent Residents

GeneBench-Pro doesn’t have a direct local footprint, but the broader direction of travel matters for anyone who uses NHS services in Kent. AI systems evaluated on benchmarks like this — covering cancer genomics, pharmacogenomics and genetic risk stratification — could eventually support more accurate diagnosis and personalised treatment through services such as NHS Kent and Medway Integrated Care Board, though any clinical use would require full regulatory approval and integration into NHS workflows. Researchers at the University of Kent working in computational biology or life sciences may also find GeneBench-Pro a useful tool for assessing AI models in their own work. For now, the benchmark is a measure of how far the technology still has to go, rather than a signal that AI-driven genomics is arriving in GP surgeries any time soon.

Source: @OpenAI

OpenAI Launches GeneBench-Pro to Test AI Judgment on Complex Biological Data Quiz

5 questions