Back

Widespread Regulatory Turnover Across Human Segmental Duplications

Morrissey, A.; Brown, A. S.; Yang, J.; Mahony, S.

2026-01-22 bioinformatics
10.64898/2026.01.21.700947 bioRxiv
Show abstract

Segmental duplications (SDs) are a pervasive feature of eukaryotic genomes, enabling genomic innovation via the duplication of genes and cis-regulatory elements. However, the regulatory mechanisms that govern these regions of the genome after duplication are not well understood. The repetitive nature of SDs pose problems for short-read assays such as ChIP-seq and DNase-seq as they create reads that map to more than one location along the genome. Moreover, the newest Telomere-to-Telomere (T2T) human genome assembly has revealed previously unknown segmental duplications. In this study, we used the Telomere-to-Telomere genome assembly along with the probabilistic allocation of multi-mapped reads by our software Allo to better understand the regulatory mechanisms governing segmental duplications in the human genome. Using 134 transcription factor ChIP-seq datasets in GM12878, we found that transcription factor binding sites within segmental duplications are rarely conserved across paralogous copies. Similarly, chromatin states defined via histone ChIP-seq datasets are also poorly conserved across segmentally duplicated regions. In contrast, paralogous genes within SDs exhibit highly correlated expression patterns. Hi-C analysis revealed that many SD paralogs have higher than expected spatial proximity, suggesting that three-dimensional genome organization may buffer transcriptional divergence following duplication despite the loss of cis-regulatory regions. Together, our results suggest that regulatory regions undergo rapid turnover following segmental duplication, but the regulatory impact may be buffered by spatial proximity within the nucleus. Tissue-specific genes and transcription factors showed greater divergence, highlighting the role of SDs in expanding specialized cellular functions.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.