Back

Augment Single-cell RNA-seq data with Surface Protein Levels using Gene set-based Deep Learning and Transfer Learning Methods

Hasib, M. M.; Zhang, T.; Zhang, J.; Gao, S.-j.; Huang, Y.

2024-05-01 bioinformatics
10.1101/2024.04.29.591655 bioRxiv
Show abstract

As scRNA-seq becomes increasingly accessible, providing a cost-efficient method to augment surface protein levels from gene expression measurements are desirable. We proposed a machine learning approach that includes a novel geneset neural network (GS-NN) that aims to learn robust and biologically meaningful features and a highly efficient transfer learning strategy to address cross-dataset differences. We conducted comprehensive experiments to show the improvements of the proposed methods. Specifically, we demonstrate that GS-NN learns more robust features to achieve better cross-subject performance than other machine learning approaches. Transfer learning further improves that of GS-NN by reducing dataset differences through highly efficient fine-tuning. The unique genesets design of GS-NN also allows identification of functions contributing to the prediction and improvement of the proposed strategy. Overall, this study reports a novel approach to robustly augment. Key PointsO_LIThe article presents a machine learning approach, Geneset Neural Network(GS-NN) to augment surface protein levels from single-cell RNA sequencing(scRNA-seq) gene expression data. C_LIO_LIThe GS-NN aims to learn robust and biologically meaningful features, and the approach includes a highly efficient transfer learning strategy to address cross-dataset differences in scRNA-seq data. C_LIO_LIComprehensive experiments demonstrate that GS-NN learns more robust features using trasfer learning techniques achieving better cross-subject performance compared to other machine learning approaches. C_LIO_LIThe unique geneset-based architecture of GS-NN allows the identification and interpretion of biological functions contributing to the prediction of cell surface protein level. C_LIO_LIGS-NNs architecture is conveniently transferrable across datasets, making it valuable tool for researchers working with diverse scRNA-seq datasets. C_LI

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.