EnhancerDetector: Enhancer Discovery from Human to Fly via Interpretable Deep Learning
Solis, L. M.; Sterling-Lentsch, G.; Halfon, M. S.; Girgis, H. Z.
Show abstract
Deciphering how enhancers encode regulatory information in DNA remains a central genomics challenge, as sequencing outpaces functional annotation. A key question is whether enhancers possess an intrinsic, sequence-based "enhancerness" distinguishing them from other regions, independent of species, cell type, or assay. Confirming its existence and learnability is both biologically fundamental and essential for scalable genome annotation. We introduce EnhancerDetector, a convolutional neural network-based framework for cross-species enhancer prediction that combines high accuracy with biological interpretability. Trained on human data, EnhancerDetector achieves strong performance across human, mouse, and fly datasets, consistently outperforming existing methods in precision and F1. It generalizes to datasets generated using diverse experimental assays. Unlike chromatin feature-based predictors requiring complex post hoc thresholding, EnhancerDetector directly outputs enhancer probability scores from short sequence windows, simplifying enhancer discovery workflows. An ensemble strategy further improves prediction reliability by reducing false positives. EnhancerDetector supports fine-tuning on new species and retains strong performance even when adapted with as few as 20,000 enhancer sequences, making it ideal for newly sequenced genomes with limited experimental data. For interpretability and visualization, we apply class activation maps to identify sequence regions predictive of enhancer activity. Experimental validation in transgenic flies confirms the predictive power of EnhancerDetector: five of six tested candidates drove reporter expression, and four exhibited expression patterns supported by prior literature. These analyses highlight distinct sequence and contextual features that confer what we term "enhancerness:" enhancer sequences possess a characteristic, identifiable signature.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DECODE: A Deep-learning Framework for Condensing Enhancers and Refining Boundaries with Large-scale Functional Assays 96%
- CENTRE: A gradient boosting algorithm for Cell-type-specific ENhancer-Target pREdiction 95%
- Gkmexplain: Fast and Accurate Interpretation of Nonlinear Gapped k-mer Support Vector Machines Using Integrated Gradients 94%
Similar papers in this journal
- Quantifying the Tissue-Specific Regulatory Information within Enhancer DNA Sequences 96%
- Accurate prediction of cis-regulatory modules reveals a prevalent regulatory genome of humans 94%
- Sequence-based chromatin activity modeling and regulatory impact prediction of genetic variants in farmed animals using deep learning 94%
Similar papers in this journal
- Genome-wide Enhancer Maps Differ Significantly in Genomic Distribution, Evolution, and Function 94%
- Reporter gene assays and chromatin-level assays define substantially non-overlapping sets of enhancer sequences 94%
- A map of cis-regulatory modules and constituent transcription factor binding sites in 80% of the mouse genome 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.