Back

MAGE: Monte Carlo method for Aberrant Gene Expression

Beltran, M.; Joh, R. I.

2024-11-11 bioinformatics
10.1101/2024.11.08.622686 bioRxiv
Show abstract

Identifying genes that are aberrantly expressed is an important first step in the diagnosis and treatment of many diseases. Conventionally, differential expression (DE) analysis is used to screen gene expression profiles to identify functionally associated genes. DE often relies on the variance and fold change in expression from individual genes, which does not consider the expression of all other genes within the profile. When the overall gene expression is skewed, DE does not capture outliers in gene expression. To address this, we have developed a non-parametric DE method based on the probability density for an entire expression profile to select genes that deviate from the global distribution between two gene expression profiles with multiple replicates. Rather than assuming a particular distribution of expression per gene, our method assumes that aberrantly expressed genes (AEGs) will exhibit expression patterns distinguishable from non-AGEs which make up the majority of the profile. Here we introduce our nonparametric method (MAGE: Monte Carlo method for aberrant gene expression) and demonstrate that MAGE can identify AEGs that are not found by conventional DE analyses. The main feature of MAGE is (1) identifying outliers based on the expression profile of all genes rather than performing DE analyses on a per-gene basis and (2) consideration of the variance in expression between two different conditions. We also compared our results with traditional DE analysis as well as density-based clustering methods. MAGE produces consistent results in a variety of conditions and performs conservatively with the addition of noise. We also applied MAGE to single-cell RNA-seq samples and demonstrated that the analysis is robust with subsampling.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.