Back

Predictive Gene Discovery with EPCY: A Density-Based Alternative to DE analysis

Audemard, E. O.; Spinella, J.-F.; Lavallee, V.-P.; Hebert, J.; Sauvageau, G.; Lemieux, S.

2025-08-11 bioinformatics
10.1101/2025.08.07.668357 bioRxiv
Show abstract

Identifying predictive genes from high-throughput data remains a key challenge in biomedical research. Most current approaches rely on statistical tests to select differentially expressed genes (DEGs), which may not align with the goal of predicting outcomes. We present EPCY, a method that ranks genes based on their predictive power using cross-validated classifiers and density estimation, without relying on null hypothesis testing. Applied to both bulk and single-cell RNA sequencing datasets, EPCY consistently outperforms benchmark DEG-based methods in selecting robust candidate genes. It also demonstrates greater stability across varying cohort sizes, enabling reproducible gene prioritization even in large, heterogeneous datasets. EPCY provides interpretable predictive scores, facilitating candidate selection aligned with downstream validation goals.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.