Back

Single-sample proteome enrichment enables missing protein recovery and phenotype association

Wong, B. J. H.; Kong, W.; Goh, W. W. B.

2021-11-15 bioinformatics
10.1101/2021.11.13.468488 bioRxiv
Show abstract

Proteomic studies characterize the protein composition of complex biological samples. Despite recent developments in mass spectrometry instrumentation and computational tools, low proteome coverage remains a challenge. To address this, we present Proteome Support Vector Enrichment (PROSE), a fast, scalable, and effective pipeline for scoring protein identifications based on gene co-expression matrices. Using a simple set of observed proteins as input, PROSE gauges the relative importance of proteins in the phenotype. The resultant enrichment scores are interpretable and stable, corresponding well to the source phenotype, thus enabling reproducible recovery of missing proteins. We further demonstrate its utility via reanalysis of the Cancer Cell Line Encyclopedia (CCLE) proteomic data, with prediction of oncogenic dependencies and identification of well-defined regulatory modules. PROSE is available as a user-friendly Python module from https://github.com/bwbio/PROSE.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.