MS2Prop: A machine learning model that directly predicts chemical properties from mass spectrometry data for novel compounds
Voronov, G.; Frandsen, A.; Bargh, B.; Healey, D.; Lightheart, R.; Kind, T.; Dorrestein, P. C.; Colluru, V.; Butler, T.
Show abstract
Mass spectrometry (MS) is a fundamental analytical tool for the study of complex molecular mixtures and in natural products drug discovery and metabolomics specifically, due to its high sensitivity, specificity, and throughput. A major challenge, however, is the lack of structurally annotated mass spectra for these applications. This deficiency is particularly acute for analyses conducted on extracts or fractions that are largely chemically undefined. This work describes the use of mass spectral data in a fundamentally different manner than structure determination; to predict properties or activities of structurally unknown compounds without the need for defined or deduced chemical structure using a machine learning (ML) model, MS2Prop. The models predictive accuracy and scalability is benchmarked against commonly used methods and its performance demonstrated in a natural products drug discovery setting. A new cheminformatic subdiscipline, quantitative spectra-activity relationships (QSpAR), using spectra rather than chemical structure as input, is proposed to describe this approach and to distinguish it from structure based quantitative methods.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Group-walk, a rigorous approach to group-wise false discovery rate analysis by target-decoy competition 94%
- Attention-based approach to predict drug-target interactions across seven target superfamilies 94%
- Identification of metabolites from tandem mass spectra with a machine learning approach utilizing structural features 94%
Similar papers in this journal
- SingleFrag: A deep learning tool for MS/MS fragment and spectral prediction and metabolite annotation 95%
- ChemEmbed: A deep learning framework for metabolite identification using enhanced MS/MS data and multidimensional molecular embeddings 93%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 93%
Similar papers in this journal
Similar papers in this journal
- Spec2Vec: Improved mass spectral similarity scoring through learning of structural relationships 95%
- Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions 93%
- Common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.