Differential expression analysis in single cell and spatial RNASeq without model assumptions
Margolin, G.; Tang, A.; Leikin, S.
Show abstract
Gene up(down)regulation findings in single cell and spatial RNASeq can be inconsistent despite remarkable progress in technology. False findings in high-quality samples raise concerns about assumptions behind widely accepted data analysis approaches. We therefore propose a weighted averaging approach for data analysis without assuming anything besides randomness of technical noise. This approach is closely related to prior work on statistics of cluster-randomized experiments. We show that weighing transcript counts based on measured noise variances and utilizing weighted rather than standard unweighted tests reduces both false positive and false negative findings. Our approach eliminates the need for parametrizing data distributions and/or rescaling transcript counts, which may cause artifacts by distorting and biasing the data. The resulting analysis is less complex and produces more consistent differential gene expression estimates.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GeneWalk identifies relevant gene functions for a biological context using network representation learning 96%
- ZetaSuite, A Computational Method for Analyzing Multi-dimensional High-throughput Data, Reveals Genes with Opposite Roles in Cancer Dependency 96%
- Conditional resampling improves calibration and sensitivity in single cell CRISPR screen analysis 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- noisyR: Enhancing biological signal in sequencing datasets by characterising random technical noise 95%
- Enhancing biological signals and detection rates in single-cell RNA-seq experiments with cDNA library equalization 95%
- Interpretable trajectory inference with single-cell Linear Adaptive Negative-binomial Expression (scLANE) testing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.