scRCA: a Siamese network-based pipeline for the annotation of cell types using imperfect single-cell RNA-seq reference data
liu, y.; Li, C.; Shen, L.-C.; he, Y.; guo, W.; Gasser, R.; Hua, H. X.; Song, J.; Jun, Y. D.
Show abstract
A critical step in the analysis of single-cell transcriptomic (scRNA-seq) data is the accurate identification and annotation of cell types. Such annotation is usually conducted by comparative analysis with known (reference) data sets - which assumes an accurate representation of cell types within the reference sample. However, this assumption is often incorrect, because factors, such as human errors in the laboratory or in silico, and methodological limitations, can ultimately lead to annotation errors in a reference dataset. As current pipelines for single-cell transcriptomic analysis do not adequately consider this challenge, there is a major demand for a computational pipeline that achieves high-quality cell type annotation using imperfect reference datasets that contain inherent errors (often referred to as "noise"). Here, we built a Siamese network-based pipeline, termed scRCA, that achieves an accurate annotation of cell types employing imperfect reference data. For researchers to decide whether to trust the scRCA annotations, an interpreter was developed to explore the factors on which the scRCA model makes its predictions. We also implemented 3 noise-robust losses-based cell type methods to improve the accuracy using imperfect dataset. Benchmarking experiments showed that scRCA outperforms the proposed noise-robust loss-based methods and methods commonly in use for cell type annotation using imperfect reference data. Importantly, we demonstrate that scRCA can overcome batch effects induced by distinctive single cell RNA-seq techniques. We anticipate that scRCA (https://github.com/LMC0705/scRCA) will serve as a practical tool for the annotation of cell types, employing a reference dataset-based approach.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Robust and Scalable Graph Neural Network for Accurate Single Cell Classification 98%
- SMNN: Batch Effect Correction for Single-cell RNA-seq data via Supervised Mutual Nearest Neighbor Detection 98%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 98%
Similar papers in this journal
- Cell-type annotation with accurate unseen cell-type identification using multiple references 99%
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 96%
- Inferring latent temporal progression and regulatory networks from cross-sectional transcriptomic data of cancer samples 96%
Similar papers in this journal
- ImmuCellAI: a unique method for comprehensive T-cell subsets abundance prediction and its application in cancer immunotherapy 95%
- Cross-species prediction of transcription factor binding by adversarial training of a novel nucleotide-level deep neural network 94%
- Information-Distilled Generative Label-Free Morphological Profiling Encodes Cellular Heterogeneity 94%
Similar papers in this journal
- Deep autoencoder for interpretable tissue-adaptive deconvolution and cell-type-specific gene analysis 98%
- scGCN: a Graph Convolutional Networks Algorithm for Knowledge Transfer in Single Cell Omics 98%
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.