Back

EnSEMBLE: a framework for enhancer-anchored pathway analysis that locks in enhancer-corroborated pathways from transcriptome sequencing data for biological validation

Zhang, L.; Gupta, A.; Wang, Y.; Sharma, R.; Lawal, B.; Hou, G.; Wang, X.-S.

2026-08-21 bioinformatics
10.64898/2026.08.17.745283 bioRxiv
Show abstract

Background Pathway discovery methods for transcriptome sequencing return tens to hundreds of redundant gene sets, and biologists often subjectively select the pathways fitting biological expectations. What is missing is not another statistical method, but a way to corroborate each candidate pathway against an independent, mechanistic line of evidence. Results We introduce EnSEMBLE (Enhancer-Set Enrichment & Mechanism-Based Linked Evidence), a tool that corroborates gene-level pathway enrichment with an orthogonal enhancer layer drawn from the same transcriptome sequencing data: active enhancers transcribe enhancer RNAs already present in standard RNA-seq, so a pathway's regulatory state can be scored from the very run that produced the gene-level signal, at no added cost. EnSEMBLE pairs pathway enrichments with Enhancer-Program Enrichment Analysis (EPEA), collapses redundant gene sets into process-level Themes, and retains only those that a concordant enhancer program corroborates. This dual-evidence requirement reduced reported signatures by >97% (hundreds of gene sets to 3-18 claims) across four datasets spanning cancer perturbations and iPSC-to-neuron differentiation. Surviving claims recovered expected biology--mesenchymal-program collapse upon SNAI1 knockout, regulatory convergence during neuronal differentiation--and named mechanisms pathway enrichments missed, including an mTOR-MYC-SPT5 elongation axis in rapamycin-treated PANC1 cells. A language AI agent performs narrative synthesis over deterministic statistics, with reproducibility enforced by temperature-zero inference and three-run consensus. We further provide enhancer over-representation analysis (eORA), mapping non-coding GWAS variants to the same programs to recover cell-type-selective trait associations. Conclusions EnSEMBLE shifts transcriptomic interpretation from enumerating possibilities to adjudicating evidence, yielding a compact, traceable set of enhancer-corroborated claims that identify the regulatory programs driving cellular change and prioritize them for experimental validation.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.