Back

MS2Prop: A machine learning model that directly predicts chemical properties from mass spectrometry data for novel compounds

Voronov, G.; Frandsen, A.; Bargh, B.; Healey, D.; Lightheart, R.; Kind, T.; Dorrestein, P. C.; Colluru, V.; Butler, T.

2022-10-10 bioinformatics
10.1101/2022.10.09.511482 bioRxiv
Show abstract

Mass spectrometry (MS) is a fundamental analytical tool for the study of complex molecular mixtures and in natural products drug discovery and metabolomics specifically, due to its high sensitivity, specificity, and throughput. A major challenge, however, is the lack of structurally annotated mass spectra for these applications. This deficiency is particularly acute for analyses conducted on extracts or fractions that are largely chemically undefined. This work describes the use of mass spectral data in a fundamentally different manner than structure determination; to predict properties or activities of structurally unknown compounds without the need for defined or deduced chemical structure using a machine learning (ML) model, MS2Prop. The models predictive accuracy and scalability is benchmarked against commonly used methods and its performance demonstrated in a natural products drug discovery setting. A new cheminformatic subdiscipline, quantitative spectra-activity relationships (QSpAR), using spectra rather than chemical structure as input, is proposed to describe this approach and to distinguish it from structure based quantitative methods.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.