Single-sample proteome enrichment enables missing protein recovery and phenotype association
Wong, B. J. H.; Kong, W.; Goh, W. W. B.
Show abstract
Proteomic studies characterize the protein composition of complex biological samples. Despite recent developments in mass spectrometry instrumentation and computational tools, low proteome coverage remains a challenge. To address this, we present Proteome Support Vector Enrichment (PROSE), a fast, scalable, and effective pipeline for scoring protein identifications based on gene co-expression matrices. Using a simple set of observed proteins as input, PROSE gauges the relative importance of proteins in the phenotype. The resultant enrichment scores are interpretable and stable, corresponding well to the source phenotype, thus enabling reproducible recovery of missing proteins. We further demonstrate its utility via reanalysis of the Cancer Cell Line Encyclopedia (CCLE) proteomic data, with prediction of oncogenic dependencies and identification of well-defined regulatory modules. PROSE is available as a user-friendly Python module from https://github.com/bwbio/PROSE.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 97%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 97%
- IceR improves proteome coverage and data completeness in global and single-cell proteomics 97%
Similar papers in this journal
Similar papers in this journal
- Dynamics of single-cell protein covariation during epithelial-mesenchymal transition 96%
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 96%
- Analysis and visualization of quantitative proteomics data using FragPipe-Analyst 96%
Similar papers in this journal
- Turnover and replication analysis by isotope labeling (TRAIL) reveals the influence of tissue context on protein and organelle lifetimes 95%
- Automated sample preparation with SP3 for low-input clinical proteomics 95%
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.