Back

Evaluating agentic AI for biological discovery in autonomous and copilot settings

Johri, S.; Pimenta, E. M.; Yates, J.; Fu, J.; Bao, E. L.; Jun, H.; Reardon, B.; Bacot, S.; Shady, M.; Fu, D.; Mei, W.; Camp, S. Y.; Park, J.; Van Allen, E.

2026-06-09 cancer biology
10.64898/2026.06.04.729919 bioRxiv
Show abstract

Advances in large language models (LLMs)-based artificial intelligence (AI) agents have improved their ability to execute structured analytical workflows, including standard bioinformatic pipelines for biological discovery. However, computational biology rarely consists of deterministic pipeline execution alone. Biological datasets are heterogeneous and noisy, and meaningful discovery often requires open-ended hypothesis generation and iterative reasoning over multimodal evidence. These challenges are particularly evident in multi-omic studies, where paired molecular modalities and heterogeneous clinical contexts create both opportunities and obstacles for discovery. The extent to which emerging agentic AI systems can support or automate this mode of scientific discovery remains poorly understood. Here, we systematically evaluated the capabilities and limitations of agentic AI for biological discovery using multi-omic single cell datasets spanning 11 cancer types. We developed the Multistep Multimodal Multiomic Agentic (M3A) Framework to support LLM-driven reasoning over persistent multimodal data states and to capture agentic reasoning behavior in autonomous and human-AI copilot settings. Using this framework, we assessed AI agents across complementary tasks, including autonomous cell-type annotation, generation of falsifiable biological hypotheses from gene programs, and copilot experiments testing the effect of human involvement and domain expertise. We found that current AI agents are effective at broad, systemic exploration of complex data, whereas domain experts remain critical for methodological guidance and biological synthesis across analyses. Together, our results delineate the current potential and boundaries of agentic AI in computational biology, and establish a framework for evaluating AI systems designed to support biological discovery.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 2%
15.2%
2
Patterns
78 papers in training set
Top 0.1%
11.1%
3
Bioinformatics Advances
203 papers in training set
Top 0.2%
9.9%
4
Scientific Reports
3612 papers in training set
Top 23%
4.4%
5
PLOS ONE
5266 papers in training set
Top 32%
4.4%
6
npj Systems Biology and Applications
125 papers in training set
Top 0.4%
4.1%
7
Bioinformatics
1204 papers in training set
Top 5%
3.5%
50% of probability mass above
8
Nature Communications
5641 papers in training set
Top 37%
2.8%
9
Genome Biology
637 papers in training set
Top 4%
2.8%
10
eLife
5828 papers in training set
Top 44%
2.1%
11
Cell Systems
201 papers in training set
Top 2%
2.1%
12
iScience
1154 papers in training set
Top 16%
1.7%
13
Frontiers in Bioinformatics
49 papers in training set
Top 0.6%
1.4%
14
Molecular Systems Biology
162 papers in training set
Top 2%
1.4%
15
BioData Mining
22 papers in training set
Top 0.4%
1.4%
16
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.4%
17
PeerJ
308 papers in training set
Top 7%
1.4%
18
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 32%
1.3%
19
npj Digital Medicine
118 papers in training set
Top 3%
1.1%
20
GigaScience
212 papers in training set
Top 3%
1.1%
21
BMC Genomics
406 papers in training set
Top 6%
1.1%
22
Nature Methods
385 papers in training set
Top 5%
1.1%
23
Molecular Biology of the Cell
311 papers in training set
Top 3%
1.0%
24
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
25
BMC Bioinformatics
457 papers in training set
Top 6%
0.9%
26
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.9%
27
Frontiers in Genetics
230 papers in training set
Top 5%
0.9%
28
PLOS Biology
486 papers in training set
Top 11%
0.9%
29
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.7%
0.9%
30
PNAS Nexus
159 papers in training set
Top 4%
0.6%