Cluster decomposition-based anomaly detection for rare cell identification in single-cell expression data
Xu, Y.; Wang, S.; Li, H.-D.; Feng, Q.; Li, Y.; Wang, J.
Show abstract
Single-cell RNA sequencing (scRNA-seq) technologies have been widely used to characterize cellular landscapes in complex tissues. Large-scale single-cell transcriptomics holds great potential for identifying rare cell types critical to the pathogenesis of diseases and biological processes. Existing methods for identifying rare cell types often rely on one-time clustering using partial or global gene expression. However, these rare cell types may be overlooked in the initial clustering step, making them difficult to distinguish. In this paper, we propose a Cluster decomposition-based Anomaly Detection method (scCAD), which iteratively decomposes clusters based on the most differential signals in each cluster to effectively separate rare cell types and achieve accurate identification. We benchmark scCAD on 25 real-world scRNA-seq datasets, demonstrating its superior performance compared to 10 state-of-the-art methods. In-depth case studies across diverse datasets, including mouse airway, brain, intestine, human pancreas, immunology data, and clear cell renal cell carcinoma, showcase scCADs efficiency in identifying rare cell types in complex biological scenarios. Furthermore, scCAD can correct the annotation of rare cell types and identify immune cell subtypes associated with disease, providing new insights into disease progression.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 97%
- CosGeneGate Selects Multi-functional and Credible Biomarkers for Single-cell Analysis 97%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.