From mass spectral features to molecules in molecular networks: a novel workflow for untargeted metabolomics.
Olivier-Jimenez, D.; Bouchouireb, Z.; Ollivier, S.; Mocquard, J.; Allard, P.-M.; Bernadat, G.; Chollet-Krugler, M.; Rondeau, D.; Boustie, J.; van der Hooft, J. J. J.; Wolfender, J.-L.
Show abstract
In the context of untargeted metabolomics, molecular networking is a popular and efficient tool which organizes and simplifies mass spectrometry fragmentation data (LC-MS/MS), by clustering ions based on a cosine similarity score. However, the nature of the ion species is rarely taken into account, causing redundancy as a single compound may be present in different forms throughout the network. Taking advantage of the presence of such redundant ions, we developed a new method named MolNotator. Using the different ion species produced by a molecule during ionization (adducts, dimers, trimers, in-source fragments), a predicted molecule node (or neutral node) is created by triangulation, and ultimately computing the associated molecules calculated mass. These neutral nodes provide researchers with several advantages. Firstly, each molecule is then represented in its ionization context, connected to all produced ions and indirectly to some coeluted compounds, thereby also highlighting unexpected widely present adduct species. Secondly, the predicted neutrals serve as anchors to merge the complementary positive and negative ionization modes into a single network. Lastly, the dereplication is improved by the use of all available ions connected to the neutral nodes, and the computed molecular masses can be used for exact mass dereplication. MolNotator is available as a Python library and was validated using the lichen database spectra acquired on an Orbitrap, computing neutral molecules for >90% of the 156 molecules in the dataset. By focusing on actual molecules instead of ions, MolNotator greatly facilitates the selection of molecules of interest. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=126 SRC="FIGDIR/small/473622v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@730ecaorg.highwire.dtl.DTLVardef@1d012e1org.highwire.dtl.DTLVardef@187a1e5org.highwire.dtl.DTLVardef@195c53f_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Met-ID: An Open-Source Software for Comprehensive Annotation of Multiple On-Tissue Chemical Modifications in MALDI-MSI 98%
- MS-CleanR: A feature-filtering approach to improve annotation rate in untargeted LC-MS based metabolomics 98%
- Rapid Development of Improved Data-dependent Acquisition Strategies. 97%
Similar papers in this journal
- Public LC-Orbitrap-MS/MS Spectral Library for Metabolite Identification 98%
- How much storage precision can be lost: Guidance for near-lossless compression of untargeted metabolomics mass spectrometry data 97%
- Boosting the MS1-only proteomics with machine learning allows 2000 protein identifications in 5-minute proteome analysis 97%
Similar papers in this journal
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 97%
- ModiFinder: Tandem Mass Spectral Alignment Enables Structural Modification Site Localization 97%
- Comparison of Cosine, Modified Cosine, and Neutral Loss Based Spectrum Alignment For Discovery of Structurally Related Molecules 97%
Similar papers in this journal
- WiPP: Workflow for improved Peak Picking for Gas Chromatography-Mass Spectrometry (GC-MS) data 96%
- MS2Lipid: a lipid subclass prediction program using machine learning and curated tandem mass spectral data 96%
- Scan-Centric, Frequency-Based Method for Characterizing Peaks from Direct Injection Fourier transform Mass Spectrometry Experiments 96%
Similar papers in this journal
- PeakBot: Machine learning based chromatographic peak picking 97%
- MAFFIN: Metabolomics Sample Normalization Using Maximal Density Fold Change with High-Quality Metabolic Features and Corrected Signal Intensities 97%
- LipidMS 3.0: an R-package and a web-based tool for LC-MS/MS data processing and lipid annotation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.