Why do cell-level test methods for detecting differentially expressed genes fail in single-cell RNA-seq data?
Lee, H.; Han, B.
Show abstract
Large-scale multi-subject single-cell data have become very common. For differential gene expression (DGE) analysis of these datasets, a common practice is to choose any of the two: pseudobulk methods or cell-wise methods. However, multi-subject single-cell studies have highly heterogeneous study designs. Some studies have case/control samples to be compared and tested, and some make within-sample perturbations. In this work, we report that to prevent severely inflated errors, we must treat the two categories of study designs differently in the DGE analysis. In studies with case/control labels, pseudobulk methods work the best, and cell-wise methods produce severe inflation of type II errors. Cell-wise methods work best in studies with within-sample perturbations, and pseudobulk methods produce severe inflation of type I errors. We provide mathematical proofs to support this argument. Surprisingly, many existing studies, even published ones, often choose an inappropriate DGE method. The most common temptation is to use cell-wise methods for case/control studies, which will only make p-values falsely look significant. Our analyses and proofs warn that choosing an appropriate DGE method is not an option but a requirement.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Benchmarking Differential Abundance Analysis Methods for Correlated Microbiome Sequencing Data 94%
- FIRM: Flexible Integration of single-cell RNA-sequencing data for large-scale Multi-tissue cell atlas datasets 94%
- CosGeneGate Selects Multi-functional and Credible Biomarkers for Single-cell Analysis 94%
Similar papers in this journal
- Correspondence analysis for dimension reduction, batch integration, and visualization of single-cell RNA-seq data 95%
- Data-driven detection of subtype-specific and differentially expressed genes 95%
- Controlling for background genetic effects using polygenic scores improves the power of genome-wide association studies 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.