mzLearn, a data-driven LC/MS signal detection algorithm, enables pre-trained generative models for untargeted metabolomics
Pirhaji, L.; Eaton, J.; Jeewajee, A. K.; Zhang, M.; Morris, M.; Karasarides, M.
Show abstract
Metabolite alterations are linked to diseases, yet large-scale untargeted metabolomics remains constrained by challenges in signal detection and integration of diverse datasets for developing pre-trained generative models. Here, we introduce mzLearn, a data-driven MS1 signal-detection and alignment method that runs from mzML files without user-set parameters. Across 15 public datasets, mzLearn detects 11,442 signals on average vs 7,100 (XCMS) and 4,655 (ASARI), with higher TP (89.0% vs 77.4% vs 49.6%) and lower FP (12.5% vs 17.3% vs 18.8%), while correcting instrument drifts across large cohorts without experimental QC samples. mzLearn detected 2,736 robust metabolite signals from 22 public studies (20,548 blood samples), enabling the development of pre-trained variational autoencoder for untargeted metabolomics. Learned metabolite representations reflected demographic data and when fine-tuned on unseen renal cell carcinoma data, improved risk stratification and overall survival predictions, while feature-importance analysis (SHAP) highlighted biologically plausible lipid and carnitine signals. By producing a consistent, high-quality MS1 feature matrix at scale, mzLearn paves the way for developing pre-trained foundation models for untargeted metabolomics.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sequence-to-sequence translation from mass spectra to peptides with a transformer model 96%
- AlphaPeptDeep: A modular deep learning framework to predict peptide properties for proteomics 96%
- TidyMass2: Advancing LC-MS Untargeted Metabolomics Through Metabolite Origin Inference and Metabolic Feature-based Functional Module Analysis 96%
Similar papers in this journal
Similar papers in this journal
- Algorithmic Learning for Auto-deconvolution of GC-MS Data to Enable Molecular Networking within GNPS. 95%
- Classes for the masses: Systematic classification of unknowns using fragmentation spectra 94%
- Multiplexed Single-Molecule Epigenetic Analysis of Plasma-Isolated Nucleosomes for Cancer Diagnostics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.