Back

GHT-SELEX demonstrates unexpectedly high intrinsic sequence specificity and complex DNA binding of many human transcription factors

Jolma, A.; Hernandez-Corchado, A.; Yang, A. W.; Fathi, A.; Laverty, K. U.; Brechalov, A.; Razavi, R.; Zheng, H.; The Codebook Consortium, ; Kulakovskiy, I. V.; Najafabadi, H. S.; Hughes, T. R.

2025-10-08 genomics
10.1101/2024.11.11.618478 bioRxiv
Show abstract

Precise identification of transcription factor (TF) binding sites is a long standing challenge in human regulatory genomics: TF binding motifs are short and degenerate, while the genome is large. Motif scans, therefore, often produce excessive binding site predictions. By surveying 179 TFs across 25 families using >1,500 cyclic in vitro selection experiments with fragmented, naked, and unmodified genomic DNA - a method we term GHT-SELEX (Genomic HT-SELEX) - we find that many human TFs possess much higher sequence specificity than anticipated. Moreover, genomic binding regions from GHT-SELEX are often surprisingly similar to those obtained in vivo (i.e., ChIP-seq peaks). Contrary to conventional wisdom, we find that high specificity can also be obtained from motif scans, but performance is highly dependent on the derivation and use of the motifs, including accounting for multiple local matches. We also observe alternative engagement of multiple DNA-binding domains within the same protein: long C2H2 zinc finger proteins often utilize modular DNA recognition, engaging different subsets of their DNA-binding domain (DBD) arrays to recognize multiple types of distinct target sites, frequently evolving via internal duplication and divergence of one or more DBDs. Thus, it is common for TFs to possess sufficient intrinsic specificity to delineate a large fraction of in vivo genomic targets, independently of other cellular factors.

Published in Nature Methods (predicted rank #9) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.