Hierarchical Encoding of Regulatory Mechanisms and Expression Syntax by a foundational genomic sequence-to-function model
Li, J.
Show abstract
Deciphering the comprehensive regulatory rules encoded in the genomic sequence remains a central challenge in functional genomics, requiring a paradigm shift from descriptive annotations to a mechanistic understanding of the entire genome. Here, we introduce HERMES (Hierarchical Encoding of Regulatory Mechanisms and Expression Syntax), a framework that progressively defines a fundamental sequence vocabulary for the complex regulatory genome and parses gene expression syntax into transparent biological insights. Specifically, we trained a foundational sequence-to-function model on a massive compendium of 137,127 functional genomics profiles spanning diverse biochemical marks and cellular conditions, including DNA methylation, transcription factor binding, polymerase binding, histone marks, chromatin accessibility and RNA expression across various tissues, cell lines and cell types. By harnessing the high-fidelity sequence representations of HERMES, we established a context-specific regulatory vocabulary comprising 40 distinct sequence classes and fine-grained subclusters spanning the entire genome. We interrogated the sequence determinants of diverse promoters and tissue-specific enhancers, and further identified a housekeeping genes-specific promoter sub-class, driven by ETS, YY1 and CCAAT motifs. To further decode transcriptional regulation, HERMES was leveraged to predict cell-type-specific gene expression and the enhancer perturbation effects. Extensive evaluations confirmed that the cis-regulatory elements and their interactions were captured to predict gene expression underlying different cellular conditions, and simultaneously revealed divergent enhancer dependencies between housekeeping genes (HKGs) and highly variable genes (HVGs). Ultimately, we distilled the complex sequence-to-function model into a biologically interpretable rulebook of Enhancer-Promoter interaction grammar. Synthesizing all the insights, we propose a unified compatibility model, where HKGs utilize a strong-promoter architecture for high-output expression, while HVGs depend on a context-dependent syntax driven by compatible promoters and enhancers. In summary, HERMES bridges the gap between predictive modeling and biological mechanism, transforming sequence representations into a comprehensive functional encyclopedia and a quantitative grammar of the complex regulatory genome.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HALO: Hierarchical Causal Modeling for Single Cell Multi-Omics Data 97%
- Gapped-kmer sequence modeling robustly identifies regulatory vocabularies and distal enhancers conserved between evolutionarily distant mammals 97%
- Boosting the detection of enhancer-promoter loops via novel normalization methods for chromatin interaction data 97%
Similar papers in this journal
- CREaTor: zero-shot cis-regulatory pattern modeling with attention mechanisms 98%
- An interpretable bimodal neural network characterizes the sequence and preexisting chromatin predictors of induced TF binding 97%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 97%
Similar papers in this journal
- Iterative deep learning-design of human enhancers exploits condensed sequence grammar to achieve cell type-specificity 98%
- Multiome Perturb-seq unlocks scalable discovery of integrated perturbation effects on the transcriptome and epigenome 96%
- Conserved epigenetic regulatory logic infers genes governing cell identity 95%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 97%
- Massively parallel reporter assay-informed modeling improves prediction of context-specific enhancer-gene regulatory interactions 97%
- Uncovering topologically associating domains from three-dimensional genome maps with TADGATE 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.