Back

Comprehensive evaluation of LLM capabilities for interpretation and analysis of genome-scale metabolic models in metabolic engineering

Yeoh, J. W.; Patro, C. P. K.; Wong, L.; Poh, C. L.

2026-06-08 systems biology
10.64898/2026.06.03.730004 bioRxiv
Show abstract

Genome-scale metabolic models (GSMs) underpin pathway and strain engineering by linking genes to metabolic reactions and enabling system-level simulation of cellular fluxes and intervention effects, yet end-to-end analysis workflows remain fragmented, expert-demanding, and slow to adapt. Large language models (LLMs) could transform this landscape, lowering the barrier by explaining concepts, interpreting GSM files, and turning natural-language instructions into valid analysis code, thereby substantially mitigating the time, effort, and expertise required. However, their reliability for domain-specific tasks remains unexplored. Here, we delivered a systematic benchmark of four leading LLMs (GPT-4, Gemini, Claude, DeepSeek-R1) across four task areas central to metabolic engineering: domain knowledge, metabolic flux prediction, pathway construction, and flux optimization. For benchmarking, we introduced a standardized, rubric-based evaluation framework that uses multi-LLM automated scoring (an ensemble of LLM-as-a-judge assessments) and two distinct sets of nine task-tailored metrics (domain vs coding-focused tasks), rated on a 1-5 scale (up to 45 per task), covering scientific validity and code executability where applicable. Across tasks, we reveal consistent strengths (conceptual explanation, code synthesis) and critical failure modes (e.g., context window limitations, incorrect identifier assumptions, strain-dependent reasoning errors, and errors in domain-specific algorithms). In aggregate, DeepSeek-R1 led in domain tasks, narrowly edging GPT-4, Claude, and Gemini, demonstrating that conceptual biological logic remains highly invariant across architectures. In contrast, Gemini achieved the highest score for coding tasks, distinguished by functional execution and excelled in error handling, documentation, and readability, followed by GPT-4, Claude, and DeepSeek. We also evaluated LLM self-inspection capability by injecting subtle, consequential faults: a stoichiometric sign error causing mass imbalance and an omitted pathway reaction. We reveal that conversational "blind search" prompting completely fails to localize these network faults. Instead, robust error localization requires prompts reframed with domain-informed constraints that force the LLM to leverage tool-assisted code procedures, such as COBRApy mass-balance functions. Together, this work establishes an evidence-based baseline for LLM-enabled GSM analysis, providing actionable guidance for building reliable, automation-ready workflows for pathway and strain design. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=88 SRC="FIGDIR/small/730004v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@442a55org.highwire.dtl.DTLVardef@1374927org.highwire.dtl.DTLVardef@a3b7a8org.highwire.dtl.DTLVardef@6ea308_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 14%
12.5%
2
Cell Systems
201 papers in training set
Top 0.2%
11.7%
3
Molecular Systems Biology
162 papers in training set
Top 0.1%
9.7%
4
npj Systems Biology and Applications
125 papers in training set
Top 0.2%
6.2%
5
iScience
1154 papers in training set
Top 1%
6.2%
6
Cell Reports Methods
165 papers in training set
Top 0.4%
4.8%
50% of probability mass above
7
PLOS Computational Biology
1863 papers in training set
Top 10%
3.2%
8
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
3.2%
9
Bioinformatics
1204 papers in training set
Top 5%
3.2%
10
Scientific Reports
3612 papers in training set
Top 37%
3.1%
11
npj Digital Medicine
118 papers in training set
Top 2%
2.6%
12
Advanced Science
286 papers in training set
Top 3%
2.4%
13
Genome Biology
637 papers in training set
Top 5%
2.1%
14
ACS Synthetic Biology
287 papers in training set
Top 1%
2.1%
15
Bioinformatics Advances
203 papers in training set
Top 3%
2.1%
16
Patterns
78 papers in training set
Top 1%
2.1%
17
mSystems
394 papers in training set
Top 5%
1.4%
18
eLife
5828 papers in training set
Top 60%
1.1%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 38%
1.0%
20
Communications Medicine
113 papers in training set
Top 4%
1.0%
21
Life Science Alliance
285 papers in training set
Top 8%
0.8%
22
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
23
Nature Biotechnology
172 papers in training set
Top 4%
0.8%
24
Nature Methods
385 papers in training set
Top 6%
0.8%
25
Genome Research
468 papers in training set
Top 7%
0.6%
26
Metabolic Engineering
75 papers in training set
Top 0.8%
0.6%
27
Imaging Neuroscience
282 papers in training set
Top 4%
0.6%
28
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 3%
0.6%
29
PLOS ONE
5266 papers in training set
Top 65%
0.6%