Back

GeneBench-Pro: Evaluating Multistage Statistical Reasoning\\in Genomics, Quantitative Biology, and Translational Biomedicine

Li, J. H.; Ho, A. J.

2026-06-30 bioinformatics
10.64898/2026.06.29.735386 bioRxiv
Show abstract

We introduce GeneBench-Pro, an expanded and improved version of GeneBench that comprises harder problems across a wider breadth of domains. GeneBench-Pro is a benchmark for AI agents performing realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine which seeks to capture the complexity of real-world problems that computational life scientists face when tasked with producing a conclusion upon which a downstream scientific or translational decision is contingent. The benchmark comprises 129 evaluations targeting quantities of direct practical relevance across 10 primary domains and 21 terminal subdomains, with a genomics-centered core. Similarly to GeneBench, each problem provides the agent with brief context, a target estimand, and minimal guidance otherwise; the agent must then navigate multiple dependent decision points; i.e., substantive inferential forks where a plausible wrong choice changes the downstream analysis, to identify and execute the correct analysis workflow and arrive at the correct answer. Relative to GeneBench, GeneBench-Pro adds 29 new problems, drops three, and introduces significantly redesigned versions of 54 of the remaining 100 overlapping problems. 82 of the 129 problems were reviewed by external domain experts, whose findings led to prompt/data modifications and redesign of those problems whose targets were not sufficiently identifiable. Ten externally reviewed problems are released publicly, 50 held-out problems were provided to Artificial Analysis for independent third-party model benchmarking, and the remainder are retained as an internal holdout. In evaluations over the full 129-problem suite, GPT-5.6 Sol reaches an eval-level pass rate of 28.7% at the max reasoning level, and GPT-5.6 Sol Pro reaches 31.5% in separately reported GPT Pro runs. GPT-5.5 reaches 12.0%, GPT-5.4 reaches 8.9%, and the strongest non-GPT baseline, Claude Opus 4.8, reaches 16.0%. As with GeneBench, models often complete substantial portions of the workflow but exhibit a consistent gap between noticing and acting by identifying local diagnostic signals but failing to propagate the implications to the corresponding analysis decision. As a result, models often select wrong estimators or persist on initially plausible but incorrect analysis paths. GeneBench-Pro therefore measures an emerging capability of long-horizon biological reasoning that remains unreliable.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 3%
9.8%
2
PLOS Computational Biology
1863 papers in training set
Top 5%
7.3%
3
Bioinformatics Advances
203 papers in training set
Top 0.6%
6.7%
4
Cell Systems
201 papers in training set
Top 0.6%
6.2%
5
Nature Communications
5641 papers in training set
Top 25%
6.2%
6
Scientific Reports
3612 papers in training set
Top 19%
5.1%
7
PLOS ONE
5266 papers in training set
Top 32%
4.4%
8
GigaScience
212 papers in training set
Top 0.8%
4.3%
50% of probability mass above
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
10
BMC Bioinformatics
457 papers in training set
Top 3%
2.4%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
2.0%
12
Nature Methods
385 papers in training set
Top 4%
1.9%
13
Genome Biology
637 papers in training set
Top 5%
1.9%
14
Nature
645 papers in training set
Top 6%
1.7%
15
BioData Mining
22 papers in training set
Top 0.3%
1.7%
16
Frontiers in Genetics
230 papers in training set
Top 3%
1.7%
17
Patterns
78 papers in training set
Top 1%
1.7%
18
iScience
1154 papers in training set
Top 21%
1.4%
19
BMC Genomics
406 papers in training set
Top 5%
1.4%
20
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.4%
1.1%
21
Journal of Computational Biology
48 papers in training set
Top 0.8%
1.1%
22
eLife
5828 papers in training set
Top 60%
1.1%
23
GENETICS
483 papers in training set
Top 4%
1.0%
24
Scientific Data
209 papers in training set
Top 2%
0.9%
25
Genetics in Medicine
78 papers in training set
Top 1.0%
0.8%
26
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
0.8%
27
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 1%
0.8%
28
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.7%
0.8%
29
Nature Machine Intelligence
70 papers in training set
Top 3%
0.6%
30
ACS Synthetic Biology
287 papers in training set
Top 3%
0.6%