Superscan: Supervised Single-Cell Annotation
Shasha, C.; Tian, Y.; Mair, F.; Miller, H. E. R.; Gottardo, R.
Show abstract
Automated cell type annotation of single-cell RNA-seq data has the potential to significantly improve and streamline single cell data analysis, facilitating comparisons and meta-analyses. However, many of the current state-of-the-art techniques suffer from limitations, such as reliance on a single reference dataset or marker gene set, or excessive run times for large datasets. Acquiring high-quality labeled data to use as a reference can be challenging. With CITE-seq, surface protein expression of cells can be directly measured in addition to the RNA expression, facilitating cell type annotation. Here, we compiled and annotated a collection of 16 publicly available CITE-seq datasets. This data was then used as training data to develop Superscan, a supervised machine learning-based prediction model. Using our 16 reference datasets, we benchmarked Superscan and showed that it performs better in terms of both accuracy and speed when compared to other state-of-the-art cell annotation methods. Superscan is pre-trained on a collection of primarily PBMC immune datasets; however, additional data and cell types can be easily added to the training data for further improvement. Finally, we used Superscan to reanalyze a previously published dataset, demonstrating its applicability even when the dataset includes cell types that are missing from the training set.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The impacts of active and self-supervised learning on efficient annotation of single-cell expression data 96%
- Unveiling the Power of High-Dimensional Cytometry Data with cyCONDOR 95%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 95%
Similar papers in this journal
Similar papers in this journal
- Non-linear Archetypal Analysis of Single-cell RNA-seq Data by Deep Autoencoders 95%
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 95%
- STREAK: A Supervised Cell Surface Receptor Abundance Estimation Strategy for Single Cell RNA-Sequencing Data using Feature Selection and Thresholded Gene Set Scoring 94%
Similar papers in this journal
- Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples 96%
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 95%
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.