DNA sequence quantitatively encodes CTCF-binding affinity at genome scale
Yin, Z.; Wang, Y.; Hou, G.; Chen, Y.; Feng, H.; Zhu, J.; Xie, Y.; Wang, R.; Zhang, Z.; Wu, T.; Huang, H.; Xie, S.; Wang, W.; Gu, W.; Wu, Q.; Shen, N.; Guo, Y.
Show abstract
CTCF is a central architectural protein that shapes 3D genome organization through sequence-specific DNA binding, but how DNA sequence quantitatively determines CTCF-binding strength remains poorly understood. Progress has been hampered by the lack of large, high-quality measurements of binding affinity. Here, we experimentally determine in vitro CTCF-binding affinity for 276,765 DNA sequences derived from the human genome, generating a comprehensive quantitative landscape of CTCF-DNA interactions. Leveraging these data, we develop DeepCTCF, a deep learning model that predicts CTCF-binding strength directly from DNA sequence and enables quantitative interpretation of CTCF motif grammar. Using this framework, we systematically dissect how specific sequence features modulate CTCF-binding affinity and generate quantitative predictions for disease-associated variants that alter CTCF binding. Together, this study defines general principles by which DNA sequence encodes CTCF-binding affinity and provides a quantitative framework for interpreting regulatory sequence variation. TeaserMapping how DNA encodes CTCF binding reveals quantitative rules across over 1{per thousand} of the human genome.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cancer-associated DNA Hypermethylation of Polycomb Targets Requires DNMT3A Dual Recognition of Histone H2AK119 Ubiquitination and the Nucleosome Acidic Patch 97%
- Substrate deformation regulates DRM2-mediated DNA methylation in plants 97%
- Pronounced sequence specificity of the TET enzyme catalytic domain guides its cellular function 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.