Back

Machine Learning Predictions Surpass Individual mRNAs as a Proxy of Single-cell Protein Expression

Fisher, J.; Wood, O.; Bullers, S.; Murray, L.; Li, L.; Jackson-Wood, M.

2024-12-16 bioinformatics
10.1101/2024.12.11.627925 bioRxiv
Show abstract

BackgroundExpansive repositories of single-cell RNA-seq data are now available. These data are often analysed assuming that mRNA abundance reflects the expression of their cognate proteins. However, post-transcriptional and translational regulation make mRNA an inadequate proxy for protein. High sparsity in low abundance mRNAs from single-cell transcriptomics data further complicates the extrapolation of protein expression levels. Although methods for single-cell surface protein quantification exist, they incur additional technical steps at greater expense and have yet to see wide-spread adoption. Computational approaches for protein imputation from scRNAseq data have been published, which learn transcriptome-wide patterns that predict protein expression. These models can then be applied to infer surface protein expression on RNA-seq only data, to increase the utility of existing data repositories. ResultsWe tested 8 such methods and compared the accuracy of predictions between approaches, and against cognate mRNAs as a direct proxy. Predictions from the trained models outperformed the use of mRNA expression as a proxy. We identify notable cases of cell surface proteins with very poor correlation with their mRNA that were predicted very successfully by imputation using the whole transcriptome. We find cell type signatures are a major determinant of predicted protein levels and, as such, prediction methods require representative training data. ConclusionsThese results reiterate that mRNA level is not a reliable predictor of cell surface protein expression, and that whole-transcriptome informed imputation can improve protein estimations given appropriately trained models.

Published in Genome Biology (predicted rank #1) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.