LN's t-Test: A Principled Approach to t-Testing in Single-Cell RNA Sequencing
Kviman, O.; Jun, S.-H.; Lagergren, J.
Show abstract
Single-cell RNA sequencing (scRNA-seq) has revolutionized the study of cellular heterogeneity, yet differential gene expression (DGE) analysis remains hindered by inconsistencies in log fold change (LFC) estimation. Existing methods, such as those implemented in Scanpy and Seurat, rely on log-transformed count data with a pseudocount, introducing a bias that compromises the reliability of the statistical inference. In this work, we propose LNs t-test, a novel approach to DGE testing that circumvents these biases by employing a log-Normal (LN) distribution-based LFC estimator. Our method jointly estimates the probability of non-zero expression and the mean of positive expression values, enabling an asymptotically unbiased and normally distributed LFC estimator with corresponding confidence intervals. Through extensive simulation studies, we demonstrate that LNs t-test outperforms competing methods by reducing false discovery rates and providing more accurate effect size estimates. Notably, we leverage stochastic ordering theory to explain why conventional t-tests systematically mis-classify non-differentially expressed genes under realistic variance conditions. Our approach offers a theoretically principled and computationally efficient alternative for DGE analysis in scRNA-seq, with implications for improving the reliability and interpretability of single-cell transcriptomics studies. Code that implements the results is available on GitHub: https://github.com/okviman/DE-ZILN.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Addressing Erroneous Scale Assumptions in Microbe and Gene Set Enrichment Analysis 96%
- Estimating Transfer Entropy in Continuous Time Between Neural Spike Trains or Other Event-Based Data 96%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
Similar papers in this journal
- The statistics of k-mers from a sequence undergoing a simple mutation process without spurious matches 96%
- MCMC-CE: A Novel and Efficient Algorithm for Estimating Small Right-Tail Probabilities of Quadratic Forms with Applications in Genomics 96%
- NetMix: A network-structured mixture model for reduced-bias estimation of altered subnetworks 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.