Back

BioML-bench: Evaluation of AI Agents for End-to-End Biomedical ML

Miller, H. E.; Greenig, M.; Tenmann, B.; Wang, B.

2025-09-04 bioinformatics
10.1101/2025.09.01.673319 bioRxiv
Show abstract

Large language model (LLM) agents hold promise for accelerating biomedical research and development (R&D). Several biomedical agents have recently been proposed, but their evaluation has largely been restricted to question answering (e.g., LAB-Bench) or narrow bioinformatics tasks. Presently, there remains a lack of benchmarks evaluating agent capability in multi-step data analysis workflows or in solving the machine learning (ML) challenges central to AI-driven therapeutics development, such as perturbation response modeling or drug toxicity prediction. We introduce BioML-bench, the first benchmarking suite for evaluating AI agents on end-to-end biomedical ML tasks. BioML-bench spans four domains (protein engineering, single-cell omics, biomedical imaging, and drug discovery) with tasks that require agents to parse a task description, build a pipeline, implement models, and submit predictions graded by established metrics (e.g., AUROC, Spearman). We evaluate four open-source agents: two biomedical specialists (STELLA, Biomni) and two generalists (AIDE, MLAgentBench). On average, agents underperform relative to human baselines, and biomedical specialization does not confer a consistent advantage. We also found that agents which employed more diverse ML strategies more often tended to score highest, suggesting that architecture and scaffolding may be stronger determinants of performance. These findings underscore both the potential and current limits of agentic systems for biomedical ML, and highlight the need for systematic, reproducible evaluations. BioML-bench is provided open-source at github.com/science-machine/biomlbench.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Bioinformatics Advances
203 papers in training set
Top 0.1%
23.2%
2
Bioinformatics
1204 papers in training set
Top 1%
19.2%
3
Briefings in Bioinformatics
354 papers in training set
Top 0.5%
10.0%
50% of probability mass above
4
GigaScience
212 papers in training set
Top 0.9%
4.2%
5
PLOS Computational Biology
1863 papers in training set
Top 9%
3.4%
6
Patterns
78 papers in training set
Top 0.8%
2.5%
7
iScience
1154 papers in training set
Top 10%
2.5%
8
Genome Biology
637 papers in training set
Top 4%
2.5%
9
Scientific Reports
3612 papers in training set
Top 41%
2.5%
10
BMC Bioinformatics
457 papers in training set
Top 4%
2.0%
11
PLOS ONE
5266 papers in training set
Top 46%
2.0%
12
Nature Communications
5641 papers in training set
Top 44%
1.8%
13
Nature Methods
385 papers in training set
Top 4%
1.7%
14
Cell Systems
201 papers in training set
Top 3%
1.5%
15
Frontiers in Bioinformatics
49 papers in training set
Top 0.6%
1.5%
16
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
17
Frontiers in Genetics
230 papers in training set
Top 5%
1.0%
18
Molecular Biology of the Cell
311 papers in training set
Top 4%
0.6%
19
eLife
5828 papers in training set
Top 67%
0.6%
20
Genome Research
468 papers in training set
Top 7%
0.6%
21
Protein Science
246 papers in training set
Top 4%
0.6%
22
BMC Genomics
406 papers in training set
Top 10%
0.5%
23
Nature Machine Intelligence
70 papers in training set
Top 3%
0.5%
24
International Journal of Molecular Sciences
494 papers in training set
Top 19%
0.5%