STEVE: Single-cell Transcriptomics Expression Visualization and Evaluation
Torbenson, E. J.; Ma, X.; Lin, J.-R.; Garry, D.; Jameson, S. C.; Zhang, Z.; Niedernhofer, L. J.; Zhang, L.; Li, M.; Dong, X.
Show abstract
Single-cell RNA sequencing (scRNA-seq) has become a key technology for characterizing cell-type heterogeneity in complex tissues. However, its utility depends on accurate and reproducible cell-type annotation, which remains a major analytical challenge. Although hundreds of computational tools have been developed for automated annotation, there is currently no systematic framework to evaluate annotation robustness in a dataset-specific manner or within the context of complete analytical pipelines. Here, we present STEVE (Single-cell Transcriptomics Expression Visualization and Evaluation), a quantitative framework designed to assess the accuracy, robustness, and reproducibility of cell-type annotation in scRNA-seq studies. STEVE implements three complementary in silico evaluation modules: (i) Subsampling Evaluation to quantify annotation stability under varying reference sizes and data partitions; (ii) Novel Cell Evaluation to assess the ability to detect previously unseen cell types; and (iii) Annotation Benchmarking to compare alternative annotation tools against ground-truth labels. In addition, STEVE includes a Reference Transfer Annotation module that enables cross-dataset cell-type mapping using external reference datasets. All modules are built upon a unified probabilistic framework that provides consistent confidence estimation across evaluation scenarios. We evaluated STEVE across four independent scRNA-seq datasets with experimentally defined or expert-curated cell-type labels. Our results show that annotation robustness is strongly influenced by the annotation method, biological separability, dataset complexity, and batch effects. STEVE provides a practical framework for quantifying annotation uncertainty and improving reproducibility in single-cell transcriptomic analyses. STEVE is freely available at GitHub (https://github.com/XiaoDongLab/STEVE).
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data 97%
- Beyond benchmarking: towards predictive models of dataset-specific single-cell RNA-seq pipeline performance 97%
- A systematic evaluation of highly variable gene selection methods for single-cell RNA-sequencing 97%
Similar papers in this journal
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 97%
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 96%
- A comprehensive comparison on cell type composition inference for spatial transcriptomics data 96%
Similar papers in this journal
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 96%
- Flexible comparison of batch correction methods for single-cell RNA-seq using BatchBench 96%
- MarcoPolo: a clustering-free approach to the exploration of differentially expressed genes along with group information in single-cell RNA-seq data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.