Learning from All Views: A Multiview Contrastive Framework for Metabolite Annotation
Zhou Chen, Y.; Hassoun, S.
Show abstract
Metabolomics, enabled by high-throughput mass spectrometry, promises to advance our understanding of cellular biochemistry and guide new discoveries in disease mechanisms, drug development, and personalized medicine. However, as the assignment of molecular structures to measured spectra is challenging, annotation rates remain low and hinder potential advancements. We present MultiView Projection (MVP), a novel framework for learning a joint embedding space between molecules and spectra by leveraging multiple data views: molecular graphs, molecular fingerprints, spectra, and consensus spectra. MVP builds on contrastive multiview learning to capture mutual information across views, leading to more robust and generalizable representations for spectral annotation. Unlike prior approaches that consider multiple views via concatenation or as targets of auxiliary tasks, MVP learns from all views jointly, resulting in improved molecular candidate ranking. Notably, MVP supports annotation using either individual spectra or consensus spectra, enabling flexible use of multiple measurements. On the MassSpecGym benchmark, we show that annotation using query consensus spectra significantly outperforms rank aggregation strategies based on constituent spectrum annotation. Using the consensus spectrum view, MVP achieves 35.99% and 13.96% rank@1 when retrieving candidates by mass and formula, respectively. When ranking using individual spectra, MVP demonstrates performance that is superior to or on par with existing methods, achieving 26.37% and 11.10% rank@1 for candidates by mass and formula, respectively. MVP offers a flexible, extensible foundation for learning from multiple molecule/spectra data views. For Table of Contents Only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/688047v1_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@f6c219org.highwire.dtl.DTLVardef@40f8d4org.highwire.dtl.DTLVardef@19031b6org.highwire.dtl.DTLVardef@1afa58d_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- 3DMolMS: Prediction of Tandem Mass Spectra from Three Dimensional Molecular Conformations 94%
- Target-Decoy MineR for determining the biological relevance of variables in noisy data sets 93%
- MSModDetector: A Tool for Detecting Mass Shifts and Post-Translational Modifications in Individual Ion Mass Spectrometry Data 93%
Similar papers in this journal
- MolDiscovery: Learning Mass Spectrometry Fragmentation of Small Molecules 97%
- FIDDLE: a deep learning method for chemical formulas prediction from tandem mass spectra 96%
- TidyMass2: Advancing LC-MS Untargeted Metabolomics Through Metabolite Origin Inference and Metabolic Feature-based Functional Module Analysis 94%
Similar papers in this journal
- Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data 97%
- Annotating metabolite mass spectra with domain-inspired chemical formula transformers 95%
- Deep Learning Prediction of Glycopeptide Tandem Mass Spectra Powers Glycoproteomics 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.