Back

Quantifying the contribution of DNA conformational flexibility to transcription factor binding on nucleosomal DNA uncovers indirect readout across diverse TF families

Dey, U.; Martinez, G. S.; Kumar, R.; Yella, V. R.; Kumar, A.

2026-06-06 bioinformatics
10.1101/2025.05.21.655105 bioRxiv
Show abstract

BackgroundEukaryotic gene regulation depends on transcription factors (TFs) recognizing short DNA motifs within chromatin. Many of these motifs lie within nucleosomes, where DNA is sharply bent, rotationally phased, and constrained by histone-DNA contacts. Yet only a subset is occupied in any cellular context. Motif identity alone, therefore, cannot fully explain selective TF engagement with nucleosomal DNA. We asked whether sequence-derived DNA conformational flexibility provides an interpretable representation of sequence context relevant to TF recognition on nucleosomes. ResultsWe compiled five DNA flexibility descriptors in the Python package DNAflexpy, representing bendability, torsional deformation, backbone conformational variability, and stiffness. We built quantitative models of TF binding affinity across 226 datasets from a high-throughput in vitro TF-nucleosome binding assay. Flexibility-augmented models improved prediction over mononucleotide baselines in most datasets, with smaller but reproducible gains over trinucleotide baselines. The gains were not uniform: they varied across TF families and were concordant with DNA shape-fluctuation features, suggesting that DNAflexpy descriptors capture a sequence-encoded structural signal. In PIONEAR-seq data, model performance generalized across nucleosomal templates in a TF- and sequence-dependent manner. Beyond prediction, position-resolved flexibility footprints revealed deformation signatures at cognate motifs and flanking regions across diverse TF families. For SOX11, model-derived footprints aligned with DNA shape fluctuations from nanosecond-to-microsecond molecular dynamics trajectories of SOX11-bound nucleosomes, consistent with independently observed DNA conformational dynamics and bound-state stabilization. The in vivo data showed a similar but more context-dependent pattern. OCT4 occupancy tended to correlate with local flexibility, whereas GATA3-pioneered regions showed flexibility coupled with altered rotational positioning of cognate motifs. Flexibility-augmented classifiers further improved discrimination of occupied nucleosomal motifs across ENCODE datasets. Torsional flexibility features, particularly twist dispersion and trx, were most informative for classification. ConclusionsSequence-derived DNA conformational flexibility provides a quantitative and interpretable representation of sequence context in TF recognition on nucleosomes. By augmenting sequence with structural information, these models help quantify and interpret an indirect-readout contribution in which DNA deformation tendencies may complement motif sequence and DNA shape. This framework may help explain why only selected motif instances are engaged in chromatin, without treating flexibility as independent of primary sequence.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nucleic Acids Research
1281 papers in training set
Top 0.5%
22.5%
2
Genome Research
468 papers in training set
Top 0.2%
15.2%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.3%
6.8%
4
Bioinformatics
1204 papers in training set
Top 4%
5.5%
50% of probability mass above
5
eLife
5828 papers in training set
Top 25%
4.9%
6
Molecular Biology and Evolution
542 papers in training set
Top 2%
4.4%
7
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
4.1%
8
Genome Biology
637 papers in training set
Top 3%
3.5%
9
PLOS Computational Biology
1863 papers in training set
Top 10%
3.2%
10
Nature Communications
5641 papers in training set
Top 45%
1.7%
11
Journal of Molecular Biology
232 papers in training set
Top 2%
1.7%
12
BMC Genomics
406 papers in training set
Top 5%
1.7%
13
Epigenetics
50 papers in training set
Top 0.4%
1.3%
14
Life Science Alliance
285 papers in training set
Top 5%
1.1%
15
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
16
Scientific Reports
3612 papers in training set
Top 69%
1.0%
17
Epigenetics & Chromatin
42 papers in training set
Top 0.5%
1.0%
18
iScience
1154 papers in training set
Top 31%
1.0%
19
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.9%
20
Frontiers in Genetics
230 papers in training set
Top 5%
0.9%
21
Nucleus
12 papers in training set
Top 0.1%
0.9%
22
GigaScience
212 papers in training set
Top 5%
0.6%
23
GENETICS
483 papers in training set
Top 5%
0.6%
24
Cell Reports
1498 papers in training set
Top 29%
0.6%
25
RNA
189 papers in training set
Top 1%
0.6%
26
Bioinformatics Advances
203 papers in training set
Top 5%
0.6%