Probabilistic index models for testing differential expression in single cell RNA sequencing data
Assefa, A. T.; Vandesompele, J.; Thas, O.
Show abstract
Single-cell RNA sequencing (scRNA-seq) technologies profile gene expression patterns in individual cells. It is often of interest to test for differential expression (DE) between conditions, e.g. treatment vs control or between cell types. Simulation studies have shown that non-parametric tests, such as the Wilcoxon-rank sum test, can robustly detect significant DE, with better performance than many parametric tools specifically developed for scRNA-seq data analysis. However, these rank tests cannot be used for complex experimental designs involving multiple groups, multiple factors and confounding variables. Further, rank based tests do not provide an interpretable measure of the effect size. We propose a semi-parametric approach based on probabilistic index models (PIM) that form a flexible class of models that generalize classical rank tests. Our method does not rely on strong distributional assumptions and it allows accounting for confounding factors. Moreover, it allows for the estimation of the effect size in terms of a probabilistic index. Real data analysis demonstrate that PIM is capable of identifying biologically meaningful DE. Our simulation studies also show that DE tests succeed well in controlling the false discovery rate at its nominal level, while maintaining good sensitivity as compared to competing methods.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Markov Random Field Model for Network-based Differential Expression Analysis of Single-cell RNA-seq Data 97%
- Benchmarking imputation methods for network inference using a novel method of synthetic scRNA-seq data generation 96%
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 96%
Similar papers in this journal
Similar papers in this journal
- Detection of genes with differential expression dispersion unravels the role of autophagy in cancer progression 97%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 97%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.