Back

SIEVE: One-stop differential expression, variability, and skewness analyses using RNA-Seq data

Li, H.; Khang, T. F.

2025-04-18 bioinformatics
10.1101/2024.04.09.588804 bioRxiv
Show abstract

MotivationRNA-Seq data analysis is commonly biased towards detecting differentially expressed genes and insufficiently conveys the complexity of gene expression changes between biological conditions. This bias arises because discrete models of RNA-Seq count data cannot fully characterize the mean, variance, and skewness of gene expression distribution using independent model parameters. A unified framework that simultaneously tests for differential expression, variability, and skewness is needed to realize the full potential of RNA-Seq data analysis in a systems biology context. ResultsWe present SIEVE, a statistical methodology that provides the desired unified framework. SIEVE embraces a compositional data analysis framework that transforms discrete RNA-Seq counts to a continuous form with a distribution that is well-fitted by a skew-normal distribution. Simulation results show that SIEVE controls the false discovery rate and probability of Type II error better than existing methods for differential expression analysis. Analysis of the Mayo RNA-Seq dataset for Alzheimers disease using SIEVE reveals that a gene set with significant expression difference in mean, standard deviation and skewness between the control and the Alzheimers disease group strongly predicts a subjects disease state. Furthermore, functional enrichment analysis shows that relying solely on differentially expressed genes detects only a segment of a much broader spectrum of biological aspects associated with Alzheimers disease. The latter aspects can only be revealed using genes that show differential variability and skewness. Thus, SIEVE enables fresh perspectives for understanding the intricate changes in gene expression that occur in complex diseases. AvailabilityThe SIEVE R package and source codes are available at https://github.com/Divo-Lee/SIEVE. Author SummaryIn RNA-Seq data analysis, the majority of current tools focus only on finding with significant differences in mean expression levels between groups, but other variations in gene expression patterns, such as the variance and even the skewness, are potentially important. We developed SIEVE, a new method that enables the simultaneous detection of genes that have significant differences in mean, variance and skewness between two biological conditions. The key innovation is a data transformation that renders the analysis tractable, making it possible to detect more complex patterns in RNA-Seq data. Applied to a large Alzheimers Disease dataset, we showed that SIEVE confirmed many previously known findings, as well as uncovered novel genes and biological pathways that were missed by traditional methods. SIEVE should be useful for researchers who study complex conditions by enabling a much richer class of biologically relevant genes to be detected and studied than was previously possible.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.