Back

Agentic AI for Structural Elucidation and Discovery of Drug Metabolites from Mass Spectrometry Data

Wang, X.; Patan, A.; Zhao, H. N.; Charron-Lamoureux, V.; Shin, Y.; Petras, D.; Hong, Y.; Bowen, B. P.; Northen, T. R.; Dorrestein, P. C.; Wang, M.

2026-06-26 bioinformatics
10.64898/2026.06.23.734138 bioRxiv
Show abstract

The majority of chemical signals detected in public metabolomics repositories remain structurally undefined. Large language models (LLMs) are probabilistic systems whose capacity to generate outputs beyond their training data, which can cause hallucinations, makes them also potentially suited to hypothesize structures for molecules that have never been described. We aimed to build a system that could harness this LLM generative capacity combined with domain specific tools/framework to constrain hallucination and produce validated discoveries. We developed a GNPS2 agentic AI system that interprets LC-MS/MS data by integrating spectral alignment, molecular formula inference, rule-based structural enumeration, machine learning-based spectrum prediction, and translates natural language hypotheses from domain experts into dynamically generated analytical workflows. We demonstrate the annotation of unknown drug metabolites from public data guided by chemical hypotheses. The agent predicted, and we experimentally confirmed, a phosphorylated hydroxyzine, an acetaminophen-p-coumaric acid ester, and identified two new oxidative ibuprofen-carnitine conjugates from public repositories. These results demonstrate that LLM-driven agentic reasoning, when combined with domain expertise, can indeed generate experimentally testable structural hypotheses for previously uncharacterized metabolites leveraging pan repository data.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.2%
9.5%
2
Nature Communications
5641 papers in training set
Top 25%
6.5%
3
Communications Chemistry
48 papers in training set
Top 0.1%
6.1%
4
Journal of Chemical Information and Modeling
238 papers in training set
Top 1.0%
5.0%
5
Nature Methods
385 papers in training set
Top 2%
4.7%
6
Bioinformatics
1204 papers in training set
Top 4%
4.7%
7
Analytical Chemistry
218 papers in training set
Top 0.8%
4.7%
8
Molecular Systems Biology
162 papers in training set
Top 0.6%
3.9%
9
Journal of the American Society for Mass Spectrometry
37 papers in training set
Top 0.2%
3.9%
10
Bioinformatics Advances
203 papers in training set
Top 2%
3.4%
50% of probability mass above
11
Cell Systems
201 papers in training set
Top 1%
3.3%
12
Metabolites
53 papers in training set
Top 0.3%
3.1%
13
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.7%
14
Journal of Proteome Research
234 papers in training set
Top 0.9%
2.6%
15
PLOS Computational Biology
1863 papers in training set
Top 11%
2.6%
16
Nature Biotechnology
172 papers in training set
Top 2%
2.4%
17
PLOS ONE
5266 papers in training set
Top 44%
2.3%
18
Patterns
78 papers in training set
Top 0.9%
2.3%
19
Journal of Cheminformatics
29 papers in training set
Top 0.3%
2.3%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.3%
21
Cell Reports Methods
165 papers in training set
Top 2%
1.6%
22
Angewandte Chemie International Edition
93 papers in training set
Top 1%
1.3%
23
Nature Chemical Biology
119 papers in training set
Top 2%
1.1%
24
Chemical Science
73 papers in training set
Top 1%
1.0%
25
iScience
1154 papers in training set
Top 37%
0.8%
26
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
27
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.8%
28
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 43%
0.8%
29
Nucleic Acids Research
1281 papers in training set
Top 16%
0.6%
30
Scientific Reports
3612 papers in training set
Top 80%
0.6%