ClusterDE: a post-clustering differential expression (DE) method robust to false-positive inflation caused by double dipping
Song, D.; Li, K.; Ge, X.; Li, J. J.
Show abstract
Double dipping is a well-known pitfall in single-cell and spatial transcriptomics data analysis: after a clustering algorithm finds clusters as putative cell types or spatial domains, statistical tests are applied to the same data to identify differentially expressed (DE) genes as potential cell-type or spatial-domain markers. Because the genes that contribute to clustering are inherently likely to be identified as DE genes, double dipping can result in false-positive cell-type or spatial-domain markers, especially when clusters are spurious, leading to ambiguously defined cell types or spatial domains. To address this challenge, we propose ClusterDE, a statistical method designed to identify post-clustering DE genes as reliable markers of cell types and spatial domains, while controlling the false discovery rate (FDR) regardless of clustering quality. The core of ClusterDE involves generating synthetic null data as an in silico negative control that contains only one cell type or spatial domain, allowing for the detection and removal of spurious discoveries caused by double dipping. We demonstrate that ClusterDE controls the FDR and identifies canonical cell-type and spatial-domain markers as top DE genes, distinguishing them from housekeeping genes. ClusterDEs ability to discover reliable markers, or the absence of such markers, can be used to determine whether two ambiguous clusters should be merged. Additionally, ClusterDE is compatible with state-of-the-art analysis pipelines like Seurat and Scanpy.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PseudotimeDE: inference of differential gene expression along cell pseudotime with well-calibrated p-values from single-cell RNA sequencing data 98%
- scINSIGHT for interpreting single-cell gene expression from biologically heterogeneous data 97%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 97%
Similar papers in this journal
Similar papers in this journal
- JIND: Joint Integration and Discrimination for Automated Single-Cell Annotation 95%
- scSampler: fast diversity-preserving subsampling of large-scale single-cell transcriptomic data 95%
- Resolving single-cell heterogeneity from hundreds of thousands of cells through sequential hybrid clustering and NMF 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.