Machine Learning Predictions Surpass Individual mRNAs as a Proxy of Single-cell Protein Expression
Fisher, J.; Wood, O.; Bullers, S.; Murray, L.; Li, L.; Jackson-Wood, M.
Show abstract
BackgroundExpansive repositories of single-cell RNA-seq data are now available. These data are often analysed assuming that mRNA abundance reflects the expression of their cognate proteins. However, post-transcriptional and translational regulation make mRNA an inadequate proxy for protein. High sparsity in low abundance mRNAs from single-cell transcriptomics data further complicates the extrapolation of protein expression levels. Although methods for single-cell surface protein quantification exist, they incur additional technical steps at greater expense and have yet to see wide-spread adoption. Computational approaches for protein imputation from scRNAseq data have been published, which learn transcriptome-wide patterns that predict protein expression. These models can then be applied to infer surface protein expression on RNA-seq only data, to increase the utility of existing data repositories. ResultsWe tested 8 such methods and compared the accuracy of predictions between approaches, and against cognate mRNAs as a direct proxy. Predictions from the trained models outperformed the use of mRNA expression as a proxy. We identify notable cases of cell surface proteins with very poor correlation with their mRNA that were predicted very successfully by imputation using the whole transcriptome. We find cell type signatures are a major determinant of predicted protein levels and, as such, prediction methods require representative training data. ConclusionsThese results reiterate that mRNA level is not a reliable predictor of cell surface protein expression, and that whole-transcriptome informed imputation can improve protein estimations given appropriately trained models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Curated Single Cell Multimodal Landmark Datasets for R/Bioconductor 95%
- STREAK: A Supervised Cell Surface Receptor Abundance Estimation Strategy for Single Cell RNA-Sequencing Data using Feature Selection and Thresholded Gene Set Scoring 94%
- Discovering differential genome sequence activity with interpretable and efficient deep learning 94%
Similar papers in this journal
- Identifying similar populations across independent single cell studies without data integration 95%
- scATAcat: Cell-type annotation for scATAC-seq data 95%
- A computational method for direct imputation of cell type-specific expression profiles and cellular compositions from bulk-tissue RNA-Seq in brain disorders 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.