scDEED: a statistical method for detecting dubious 2D single-cell embeddings
Xia, L.; Lee, C.; Li, J. J.
Show abstract
Two-dimensional (2D) embedding methods are crucial for single-cell data visualization. Popular methods such as t-SNE and UMAP are commonly used for visualizing cell clusters; however, it is well known that t-SNE and UMAPs 2D embedding might not reliably inform the similarities among cell clusters. Motivated by this challenge, we developed a statistical method, scDEED, for detecting dubious cell embeddings output by any 2D-embedding method. By calculating a reliability score for every cell embedding, scDEED identifies the cell embeddings with low reliability scores as dubious and those with high reliability scores as trustworthy. Moreover, by minimizing the number of dubious cell embeddings, scDEED provides intuitive guidance for optimizing the hyperparameters of an embedding method. Applied to multiple scRNA-seq datasets, scDEED demonstrates its effectiveness for detecting dubious cell embeddings and optimizing the hyperparameters of t-SNE and UMAP.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CellScope: High-Performance Cell Atlas Workflow with Tree-Structured Representation 97%
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 97%
- Atlas-scale single-cell multi-sample multi-condition data integration using scMerge2 97%
Similar papers in this journal
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 97%
- Learning interpretable representations of single-cell multi-omics data with multi-output Gaussian Processes 96%
- Inferring cell diversity in single cell data using consortium-scale epigenetic data as a biological anchor for cell identity 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.