Back

Differentiable Gene Set Enrichment Analysis for Pathway-Level Supervision in Transcriptomic Learning

Li, S.; Ruan, Y.; Yang, X.; Wen, Z.; Saigo, H.

2026-03-20 bioinformatics
10.64898/2026.03.18.712610 bioRxiv
Show abstract

In transcriptomics-driven drug discovery, upstream predictors of chemical-induced transcriptional profiles (CTPs) are typically trained with gene-wise objectives, whereas downstream interpretation relies on pathway-level, rank-based statistics such as Gene Set Enrichment Analysis (GSEA). This objective mismatch destabilizes pathway conclusions under prediction errors: small ranking perturbations can flip enrichment direction or distort pathway ordering. To bridge this gap, we present differentiable GSEA (dGSEA), a training-compatible surrogate that maps predicted gene-level scores to pathway enrichment with well-behaved gradients. Technically, dGSEA replaces discrete ranking operations with temperature-controlled soft sorting, smooth prefix accumulation, and differentiable extremum aggregation. Critically, to preserve the statistical semantics of classical GSEA, we introduce sign-specific robust permutation normalization (dNES) with optional{kappa} -calibration. For computational efficiency, a scalable Nystrom-window approximation (nyswin) reduces the quadratic bottleneck to near-linear complexity, enabling genome-scale evaluation. Empirically, across synthetic benchmarks and LINCS L1000 signatures, dGSEA matches classical GSEA accuracy with improved numerical stability. When incorporated as an auxiliary objective for SMILES-to-transcriptome prediction, dGSEA improves pathway-level agreement (macro correlation 0.257 [->] 0.306; sign accuracy 0.620 [->] 0.641) without compromising gene-level performance, providing a practical mechanism for pathway-aware optimization in transcriptomic prediction pipelines.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.