Back

A Chromatin-Structure-Guided Framework for Predictive and Interpretable Regulatory Genomics

Ye, B.; Du, L.; Chen, M.; Dai, Y.; Ma, A.; Liang, J.

2025-11-05 genomics
10.1101/2025.11.03.686435 bioRxiv
Show abstract

Chromatin organization shapes gene regulation by linking distal elements across megabase scales, yet most predictive genomics models still treat the genome as linear, without incorporating three-dimensional structure. Hi-C provides genome-wide chromatin conformation information, but its contact maps are population-averaged, distance-biased, and noisy, obscuring the biologically specific contacts. We present CHROME, a framework built on a self-avoiding polymer ensemble null model that identifies physically specific, non-random Hi-C contacts. By integrating these contacts into graph representations, CHROME enables efficient information transfer across spatially connected loci. It integrates sequence, chromatin accessibility, or pre-trained embeddings into a graph attention architecture to predict cell line-specific ChIP-seq profiles, consistently outperforming local encoder baselines and generalizing to an unseen cell line. The resulting graph embeddings also enhance prediction on tissue-specific eQTL and ClinVar variant pathogenicity, outperforming local sequence-based embeddings. Beyond predictive performance, CHROME provides interpretability through attention-derived neighbor-to-center contributions that reveal how spatially connected loci influence local regulatory activity over multi-megabase distances. Together, these results show that incorporating physically validated chromatin interactions enables more accurate and interpretable modeling of gene regulation and variant effects.

Published in Briefings in Bioinformatics (predicted rank #29) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.