Back

Reliable and accurate gene expression quantification with subpopulation structure-aware constraints for single-cell RNA sequencing

Tu, C.-C.; Hung, J.-H.

2022-11-09 bioinformatics
10.1101/2022.11.08.515740 bioRxiv
Show abstract

BackgroundSingle-cell RNA sequencing (scRNA-seq) analysis analyzes the type and state of individual cells by estimating the gene expression of each cell and enables researchers to study the biological phenomena that cannot be observed in bulk RNA sequencing. MotivationHowever, the current scRNA-seq quantification tools estimate the gene expression profile of each cell independently, ignoring the fact that there are multiple cell types in the scRNA-seq data and the expression level should be highly correlated with the cell type. Since scRNA-seq suffers from a low sequencing depth, the conventional strategy leads to a high proportion of missing values in the gene expression profile, obscuring the biological characteristics of cell subpopulations and further impacting the correctness of the subsequent downstream analysis. ResultsIn this study, we proposed Quasic, a novel scRNA-seq quantification pipeline which examines the potential cell subpopulation information during quantification, and uses the information to calculate the gene expression level. Using the human peripheral blood mononuclear cells and the simulated doublet dataset, we verified that Quasic not only correctly reinforced the cell signatures, but also identified the corresponding cell subpopulations and biological pathways more accurately. In addition, we also applied Quasic to the breast cancer cell line dataset (MCF-7), and successfully identified more potentially therapeutic resistant cells of which characteristics are consistent with that from previous studies. ConclusionsThe proposed pipeline can let the gene expression profile of each cell be more consistent with the corresponding subpopulation, making the biological features unique to the subpopulation more apparent and convenient for analysis. By using Quasic, researchers can effectively extract the desired cell subpopulation information from their sampled cells, enable them to perform cell subpopulation-related studies more accurately.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.