omicsGMF: a multi-tool for dimensionality reduction, batch correction and imputation applied to bulk- and single cell proteomics data.
Segers, A.; Castiglione, C.; Vanderaa, C.; Martens, L.; Risso, D.; Clement, L.
Show abstract
The unprecedented speed and sensitivity of mass spectrometry (MS) unlocked large-scale applications of proteomics and even enabled proteome profiling of single cells. However, this fast-evolving field is hindered by a lack of scalable dimensionality reduction tools that can compensate for substantial batch effects and missingness across MS runs. Therefore, we present omicsGMF, a fast, scalable, and interpretable matrix factorization method, tailored for bulk and single-cell proteomics data. Unlike current workflows that sequentially apply imputation, batch correction, and principal component analysis, omicsGMF integrates these steps into a unified framework, dramatically enhancing data processing and dimensionality reduction. Additionally, omicsGMF provides robust imputation of missing values, outperforming bespoke state-of-the-art imputation tools. We further demonstrate how this integrated approach increases statistical power to detect differentially abundant proteins in the downstream data analysis. Hence, omicsGMF is a highly scalable approach to dimensionality reduction in proteomics, that dramatically improves many important steps in proteomics data analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 96%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 96%
- Optimizing Proteomics Data Differential Expression Analysis via High-Performing Rules and Ensemble Inference 96%
Similar papers in this journal
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 97%
- mokapot: Fast and flexible semi-supervised learning for peptide detection 96%
- nf-encyclopedia: A cloud-ready pipeline for chromatogram library data-independent acquisition proteomics workflows 96%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 98%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 96%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 96%