Back

Interpreting biochemical text with language models:a machine learning framework for reaction extraction and cheminformatic validation

Lim, D.; Badrinarayanan, S.; Sterling, K. C.; Rajesh, G.; Mistry, E.; Yang, D.; Lee, M.; Hsu, K. B.; Manjrekar, M.; Areff, C.; Xie, P.; Kristanto, I. A.; Chandran, A.; Anderson, J. C.

2025-05-20 bioinformatics
10.1101/2025.05.15.654376 bioRxiv
Show abstract

Recent advancements in large language models (LLMs) offer new opportunities for automating the manual curation of biochemical reaction databases from scientific literature. In this study, we present an integrated pipeline that enhances LLM-based extraction of enzymatic reactions with machine learning and cheminformatics-informed validation. Using BRENDA-linked PubMed articles, we evaluate GPT-4s ability to extract reactions and infer missing chemical entities in textual descriptions of enzymatic reactions. Extracted reactions are converted to SMILES and InChI notations before being encoded into molecular fingerprint similarity scores and atom mapping metrics. These cheminformatics metrics are then used to train machine learning classifiers that validate GPT extractions. We employ a Positive-Unlabeled learning approach with synthetic invalid reactions to train various classifiers and assess model performances. The best classifier is then benchmarked on GPT extractions. Our findings show that GPT can accurately infer incomplete reactions and cheminformatics tools can serve as effective predictors of reaction validity. This work demonstrates a scalable framework for automated and reliable curation of enzymatic reaction databases, highlighting the potential of combining LLMs with cheminformatics and machine learning for reliable scientific knowledge extraction. Author SummaryCurating databases of biochemical reactions is a time-consuming and manual task, yet it plays a vital role in advancing research in biology and chemistry. Many scientific articles describe important enzymatic reactions, but often do so in incomplete ways--such as mentioning only the starting molecule or the enzyme, and leaving out the rest. In this work, we explore how recent advancements in artificial intelligence, specifically large language models like GPT, can help extract such information automatically from scientific literature. We show that these models can not only find reactions in text, but also infer missing parts of reactions based on the surrounding context. To make sure these inferred reactions are chemically plausible, we use computational chemistry tools that analyze the structure of the molecules involved. We then train a machine learning model to help us automatically detect which reactions are likely to be valid. This combination of tools offers a new way to speed up and improve how biochemical knowledge is extracted from the growing body of scientific literature. Our study suggests that this kind of automation could help scientists keep biological databases up to date and reduce the burden of manual data entry.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Journal of Cheminformatics
29 papers in training set
Top 0.1%
18.9%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.3%
15.4%
3
Bioinformatics
1204 papers in training set
Top 2%
13.0%
4
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.2%
8.1%
50% of probability mass above
5
Bioinformatics Advances
203 papers in training set
Top 1%
4.1%
6
Metabolites
53 papers in training set
Top 0.2%
4.1%
7
PLOS Computational Biology
1863 papers in training set
Top 9%
3.6%
8
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.8%
9
BMC Bioinformatics
457 papers in training set
Top 3%
2.5%
10
PLOS ONE
5266 papers in training set
Top 44%
2.2%
11
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
2.0%
12
ACS Omega
105 papers in training set
Top 2%
1.7%
13
Communications Chemistry
48 papers in training set
Top 0.7%
1.5%
14
Scientific Reports
3612 papers in training set
Top 59%
1.5%
15
Molecules
39 papers in training set
Top 0.8%
1.4%
16
Analytical Chemistry
218 papers in training set
Top 2%
1.2%
17
Chemical Science
73 papers in training set
Top 1%
1.2%
18
Frontiers in Bioinformatics
49 papers in training set
Top 1.0%
1.1%
19
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
20
Patterns
78 papers in training set
Top 3%
0.6%
21
iScience
1154 papers in training set
Top 38%
0.6%
22
Biomolecules
100 papers in training set
Top 3%
0.6%