Back

Non-consensus flanking sequence of hundreds of base pairs around in vivo binding sites: statistical beacons for transcription factor scanning

Faltejskova, K.; Sulc, J.; Vondrasek, J.

2025-06-01 bioinformatics
10.1101/2025.05.28.656598 bioRxiv
Show abstract

It was long suspected that for specific DNA binding by a transcription factor, the flanks of the binding motifs can play an important role. By a thorough analysis of the DNA sequence in the broad context ({+/-} 5000 bp) of in vivo binding sites (as identified in a ChIP-seq or a Cut&Tag experiment), we show that the average GC content is in most cases statistically significantly increased around the binding site in a patch spanning 1000-1500 bp. This increase was observed consistently in experiment targeting the same TF in different cell lines. The surrounding of binding sites of certain TFs like MYC display a directional alteration of dinucleotide frequencies. We attempt to explain these preferences by alteration in DNA shape features as well as by potential cooperation with other TFs. We observed differences in sequence affinity to various potential cooperating TFs between cell lines. Altogether, we propose that the observed feature distortion is indicative of a coarse scanning mechanism that helps TFs find the target binding site. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/656598v3_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@dac0a6org.highwire.dtl.DTLVardef@19de116org.highwire.dtl.DTLVardef@249b5eorg.highwire.dtl.DTLVardef@1544ce8_HPS_FORMAT_FIGEXP M_FIG C_FIG Key MessagesO_LIWe prove an increase in GC content 1000-1500 bp both upstream and downstream of the binding site. C_LIO_LIWe observed a funnel-like sequence signature of thousands of bp that is shared between most cell lines. C_LIO_LIWe propose an explanation based on helix shape and binding affinities of potential cooperating TFs, supported by crosslinking the ChIP-seq (Cut&Tag) experiments with ATAC-seq experiments. C_LI

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.