Prediction of Transcription Factor DNA Binding Affinity with High-Throughput Kd Measurements and Deep Learning
Wang, Z.; Wang, D.; Shen, K.; Luo, J.; Wang, X.; Wu, N.; Lang, Y.; Wang, X.; Ren, J.; Dong, W.; Pan, L.; Li, G.; Li, D.; Xie, C.; Zhang, Z.; Lyu, Y.; Yu, S.; Shan, L.; Zhang, N.; Yan, J.; Chen, M.; Xie, X. S.
Show abstract
Transcription factors (TFs) regulate gene expression through specific interactions with genomic DNA. While TF binding motifs from public databases describe sequence preferences, quantifying genome-wide affinity (Kd) is highly desirable for a more accurate thermodynamic description. Here, we report ivtFOODIE (in vitro FOOtprinting with DeamInasE), an assay that leverages deaminase-mediated cytosine-to-uracil conversion to measure Kd values for a given TF across accessible genomic regions from human cells. By pre-training on TF binding sites from JASPAR and fine-tuning with our ivtFOODIE data from 46 TFs representing 13 different DNA-binding domains (DBDs), we developed Seq2Kd, a deep learning model capable of predicting a TFs absolute binding affinity on DNA sequences. Seq2Kd enables de novo motif discovery of [~]500 previously uncharacterized human TFs and reveals the effects of genetic variation both in TF-coding regions and DNA-binding sites on gene expression and disease susceptibility. By correlating predicted affinity changes with the sign and magnitude of expression quantitative trait locus (eQTL) effects, we stratified TFs into activator-like and repressor-like groups. Compared to clinically benign variants, pathogenic single-nucleotide variants (SNVs) within regulatory and protein-coding regions show significantly larger predicted shifts in Kd. We provide an interactive web portal, the ENcyclopedia of Transcription-factor Interactions with Regulatory Elements (ENTIRE), which integrates the Seq2Kd model with the ivtFOODIE dataset. This resource offers thermodynamic prediction for TF-DNA interactions for functional genomics and human disease.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromatin information content landscapes inform transcription factor and DNA interactions 97%
- Gapped-kmer sequence modeling robustly identifies regulatory vocabularies and distal enhancers conserved between evolutionarily distant mammals 96%
- Unraveling the mechanisms of PAMless DNA interrogation by SpRY Cas9 96%
Similar papers in this journal
Similar papers in this journal
- Structural basis for Lamassu-based antiviral immunity and its evolution from DNA repair machinery 95%
- LinearTurboFold: Linear-Time Global Prediction of Conserved Structures for RNA Homologs with Applications to SARS-CoV-2 95%
- Immiscible proteins compete for RNA binding to order condensate layers 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.