DNA Conformational Flexibility Descriptors Improve Transcription Factor Binding Prediction Across the Protein Families
Dey, U.; Yella, V. R.; Kumar, A.
Show abstract
Precise binding of transcription factors (TFs) to specific DNA sequences is fundamental to gene regulation, yet the molecular principles underpinning TF-DNA specificity remain incompletely understood. While nucleotide sequence and DNA shape are known determinants of TF binding, the role of DNA flexibility encompassing axial, torsional, and stretching dynamics-- remains largely unexplored, particularly across diverse TF families. Here, we systematically integrate experimentally and computationally derived DNA flexibility descriptors into predictive models of TF-DNA binding specificity. Through extensive analyses of large-scale in vitro datasets from HT-SELEX, SELEX-Seq, protein binding microarrays encompassing mam-malian and Drosophila TFs, we demonstrate that flexibility-augmented models consistently outperform sequence based models, and DNA shape augmented models to an extent. These improvements are robust across diverse experimental platforms, and scale of the datasets, underscoring the importance of DNA conformational dynamics in indirect readout. Quantitative analyses of position-specific flexibility contributions reveal distinct "flexibility hotspots" within transcription factor binding sites and their flanking regions. This is exemplified by structural insights into the homeodomain TF MSX1, where localized DNA bendability directly correlates with enhanced binding affinity and precise recognition specificity. Finally, leveraging in vivo ChIP-Seq and DNase-Seq data from ENCODE, we further validate that DNA flexibility substantially enhances the identification of functional TF binding sites across various TF families and cellular contexts. Collectively, current findings substantiate DNA flexibility as a fundamental element of the cis-regulatory code and significantly advancing predictive frameworks of gene regulatory networks.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Domain adaptive neural networks improvecross-species prediction of transcription factor binding 96%
- TRACE: transcription factor footprinting using chromatin accessibility data and DNA sequence 95%
- Identifying transcription factor-bound gene activators and silencers in the chromatin accessible human genome using ATAC-STARR-seq 95%
Similar papers in this journal
- Identification of transcription factor co-binding patterns with non-negative matrix factorization 96%
- DNAffinity: A MACHINE-LEARNING APPROACH TO PREDICT DNA BINDING AFFINITIES OF TRANSCRIPTION FACTORS 95%
- Functional identification of cis-regulatory long noncoding RNAs at controlled false-discovery rates 94%
Similar papers in this journal
- AdaLiftOver: High-resolution identification of orthologous regulatory elements with adaptive liftOver 94%
- MAGGIE: leveraging genetic variation to identify DNA sequence motifs mediating transcription factor binding and function 94%
- CENTRE: A gradient boosting algorithm for Cell-type-specific ENhancer-Target pREdiction 94%
Similar papers in this journal
- Learning And Interpreting The Gene Regulatory Grammar In A Deep Learning Framework 95%
- Identifying Reproducible Transcription Regulator Coexpression Patterns with Single Cell Transcriptomics 94%
- TAMC: A deep-learning approach to predict motif-centric transcriptional factor binding activity based on ATAC-seq profile 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.