Back

SERM: a self-consistent deep learning solution forrapid and accurate gene expression recovery

Islam, M. T.; Wang, J.-Y.; Ren, H.; Li, X.; Khuzani, M. B.; Sang, S.; Yu, L.; Shen, L.; Zhao, W.; Xing, L.

2022-01-20 bioinformatics
10.1101/2022.01.18.476789 bioRxiv
Show abstract

Single cell RNA sequencing (scRNA-seq) is a promising technique to determine the states of individual cells and classify novel cell subtypes. Computationally, the processing of scRNA-seq data presents a daunting challenge because of the noisy nature and humongous size and dimensionality of the data. Compromised solution by omitting the genes with low expression is commonly taken in current scRNA-seq analysis, which leads to inaccurate gene counts. In this paper, we introduce a broadly applicable data-driven gene expression recovery framework, referred to as the self-consistent expression recovery machine (SERM), to impute the missing gene expression. Using deep learning, SERM first learns from a subset of the noisy gene expression data to estimate the underlying data distribution. SERM then recovers the overall gene expression data by imposing a self-consistency on the gene expression matrix, thus ensuring that the expression levels are similarly distributed in different parts of the matrix. We show that SERM significantly improves the accuracy of gene imputation with at least 100-fold increase in computational efficiency in comparison to the state-of-the-art techniques. Thus SERM promises to provide an urgently needed computational solution for rapid and accurate recovery of big genomic expression data. SERM is available as a web-based computational tool (https://www.analyxus.com/compute/serm) and its source codes can be found in https://github.com/xinglab-ai/self-consistent-expression-recovery-machine.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.