Back

Expanding the DNA Motif Lexicon of the Transcriptional Regulatory Code

Fan, J.; Chaudhri, V. K.; Bisht, D.; Pease, N. A.; Hall, D. R.; Gerges, P.; Yi, V. F.; Kales, S.; Ho, C.-H.; Das, J.; Ray, J. P.; Tewhey, R.; Meyer, C. A.; Sahni, N.; Singh, H.

2025-07-15 genomics
10.1101/2025.07.09.662874 bioRxiv
Show abstract

Transcriptional regulatory sequences in metazoans contain intricate combinations of transcription factor (TF) motifs. Stereospecific arrangements of simple motifs constitute composite elements (CEs) that enhance DNA-protein interaction specificity and enable combinatorial regulatory logic. Despite their importance, CEs remain underexplored. We advance CE discovery and functional characterization by developing an integrated framework that combines computational prediction, experimental testing and deep learning. The extended TF motif catalog comprises both synergistic and counteracting CEs, which are supported by evidence of TF binding in vivo and in vitro. A deep learning model GRACE trained on customized massively parallel reporter assays learns the lexicon of CEs at single-nucleotide resolution. Comparative analysis with a neural network model trained on chromatin accessibility demonstrates striking convergence and distinctions within the expanded regulatory lexicon, enabling joint predictions of motif contributions and the impact of variants on chromatin structure and transcriptional activity in diverse cellular contexts.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.