Back

FlowBench: separating planning, fault recovery and interpretation in agentic bioinformatics

Kurjan, A.; Cribbs, A. P.

2026-06-16 bioinformatics
10.64898/2026.06.12.731844 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWAgentic large language model (LLM) systems are being deployed in bioinformatics faster than they are understood, and single-metric evaluations conflate capabilities that fail independently. We introduce FlowBench, a benchmark that decomposes agentic bioinformatics performance into planning, fault recovery, biological interpretation, and end-to-end output-fidelity. Existing systems achieve high plan completeness, but their closed, single-provider designs prevent attribution of performance to scaffolding versus the underlying model. We therefore built FlowAgent, a modular, provider-agnostic framework whose components can be selectively disabled and whose backbone model can be swapped across providers on a shared harness, and used it to evaluate 23 models from three main providers. Three findings emerge. First, generating a valid workflow plan from a named toolchain is largely solved, whereas inferring an appropriate toolchain from biological intent alone is uniformly difficult regardless of model tier, compressing all models into a narrow 44-57% pass-rate band. Second, ablation shows that the dependency-structured plan and a completeness-reflection step drive performance, while adding a same-context validator-driven retry makes structural quality worse. Third, fault recovery and data-grounded interpretation remain unsolved. Models frequently propose fixes that force a clean exit while leaving the underlying data invalid, and data-grounded interpretation lags internal-knowledge recall by a consistent margin. Safety does not emerge from capability, and reasoning-tier models were among the least reliable at recognising unrecoverable faults. Once planning saturates, agent architecture and refusal calibration, not model scale, are the productive frontier. Availability and implementationFlowAgent and FlowBench are available under a GPLv3 licence at https://github.com/EnteloBio/flowagent Contactadam@entelo.bio

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
19.1%
2
Bioinformatics
1204 papers in training set
Top 2%
15.6%
3
Genome Biology
637 papers in training set
Top 1%
8.1%
4
PLOS Computational Biology
1863 papers in training set
Top 6%
5.6%
5
Nature Methods
385 papers in training set
Top 2%
5.6%
50% of probability mass above
6
NAR Genomics and Bioinformatics
242 papers in training set
Top 0.9%
4.5%
7
GigaScience
212 papers in training set
Top 0.9%
4.2%
8
Bioinformatics Advances
203 papers in training set
Top 1%
4.2%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.3%
10
Nature Communications
5641 papers in training set
Top 37%
2.8%
11
Nature Biotechnology
172 papers in training set
Top 2%
2.7%
12
Genome Research
468 papers in training set
Top 3%
2.2%
13
Nucleic Acids Research
1281 papers in training set
Top 9%
1.5%
14
Cell Reports Methods
165 papers in training set
Top 2%
1.2%
15
PLOS ONE
5266 papers in training set
Top 57%
1.1%
16
Molecular Systems Biology
162 papers in training set
Top 3%
0.9%
17
BMC Bioinformatics
457 papers in training set
Top 5%
0.9%
18
BMC Genomics
406 papers in training set
Top 8%
0.9%
19
Protein Science
246 papers in training set
Top 4%
0.6%
20
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.6%
21
Journal of Molecular Biology
232 papers in training set
Top 4%
0.6%
22
Nature Genetics
286 papers in training set
Top 5%
0.6%
23
Patterns
78 papers in training set
Top 3%
0.6%
24
iScience
1154 papers in training set
Top 38%
0.6%
25
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.6%