A systematic analysis of in-source fragments in LC-MS metabolomics
Chi, Y.; Mitchell, J.; Li, S.
Show abstract
The majority of features in global metabolomics from high-resolution mass spectrometry are typically not identified, referred as the "dark matter". Are these features real compounds or junk? Understanding this problem is critical to the annotation and interpretation of metabolomics data and future development of the field. Recent debates also brought attention to in-source fragments, which appear to be prevalent in spectral databases. We report here a systematic analysis of 61 representative public datasets from LC-MS metabolomics, the most common data type in biomedical studies. The results indicate that in-source fragments contribute to less than 10% of features in LC-MS metabolomics. Khipu-based pre-annotation shows that majority of abundant features have identifiable ion patterns. This suggests that the "dark matter" in LC-MS metabolomics is explainable in an abundance dependent manner; most features are from real compounds; the number of compounds is much smaller than that of features; most compounds are yet to be identified.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Met-ID: An Open-Source Software for Comprehensive Annotation of Multiple On-Tissue Chemical Modifications in MALDI-MSI 97%
- MS-CleanR: A feature-filtering approach to improve annotation rate in untargeted LC-MS based metabolomics 97%
- AutoTuner: High fidelity, robust, and rapid parameter selection for metabolomics data processing 97%
Similar papers in this journal
Similar papers in this journal
- MS2Lipid: a lipid subclass prediction program using machine learning and curated tandem mass spectral data 96%
- Scan-Centric, Frequency-Based Method for Characterizing Peaks from Direct Injection Fourier transform Mass Spectrometry Experiments 96%
- Robust Moiety Model Selection Using Mass Spectrometry Measured Isotopologues 95%
Similar papers in this journal
- Using variable data independent acquisition for capillary electrophoresis-based untargeted metabolomics 96%
- MealTime-MS: A Machine Learning-Guided Real-Time Mass SpectrometryAnalysis for Protein Identification and Efficient DynamicExclusion 96%
- Comparison of Cosine, Modified Cosine, and Neutral Loss Based Spectrum Alignment For Discovery of Structurally Related Molecules 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.