DreamAI: algorithm for the imputation of proteomics data
ma, w.; Kim, S.; Chowdhury, S.; li, z.; YANG, M.; Yoo, S.; Petralia, F.; Jacobsen, J.; Li, J. J.; Ge, X.; Li, K.; Yu, T.; Edwards, N. J.; Payne, S.; Boutros, P. C.; Rodriguez, H.; Stolovitzky, G. A.; Kang, J.; Fenyo, D.; Saez, J.; Wang, P.
Show abstract
Deep proteomics profiling using labeled LC-MS/MS experiments has been proven to be powerful to study complex diseases. However, due to the dynamic nature of the discovery mass spectrometry, the generated data contain a substantial fraction of missing values. This poses great challenges for data analyses, as many tools, especially those for high dimensional data, cannot deal with missing values directly. To address this problem, the NCI-CPTAC Proteogenomics DREAM Challenge was carried out to develop effective imputation algorithms for labeled LC-MS/MS proteomics data through crowd learning. The final resulting algorithm, DreamAI, is based on an ensemble of six different imputation methods. The imputation accuracy of DreamAI, as measured by Pearson correlation, is about 15%-50% greater than existing tools among less abundant proteins, which are more vulnerable to be missed in proteomics data sets. This new tool notably enhances data analysis capabilities in proteomics research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A machine learning strategy that leverages large datasets to boost statistical power in small-scale experiments 96%
- mokapot: Fast and flexible semi-supervised learning for peptide detection 96%
- nf-encyclopedia: A cloud-ready pipeline for chromatogram library data-independent acquisition proteomics workflows 95%
Similar papers in this journal
- Optimizing Proteomics Data Differential Expression Analysis via High-Performing Rules and Ensemble Inference 95%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 95%
- An adaptive, continuous-learning framework for clinical decision-making from proteome-wide biofluid data 95%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 98%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 96%
- Missing values are informative in label-free shotgun proteomics data: estimating the detection probability curve 96%
Similar papers in this journal
- eNODAL: an experimentally guided nutriomics data clustering method to unravel complex drug-diet interactions 94%
- Deep-learning enables proteome-scale identification of phase-separated protein candidates from immunofluorescence images 92%
- STEAM: Spatial Transcriptomics Evaluation Algorithm and Metric for clustering performance 92%