Back

Genome-wide mapping of persistent long-range correlation in the complete telomere-to-telomere human reference genome

Papatheodorou, K.; Megalovasilis, G.; Zaravinos, A.; Georgakopoulos-Soares, I.

2026-07-19 genomics
10.64898/2026.07.19.739401 bioRxiv
Show abstract

Long-range correlations (LRCs) in DNA sequences have been reported for decades, but their interpretation has been limited by incomplete representation of repetitive and structurally complex regions in earlier human genome assemblies. Here, we use the complete telomere-to-telomere (T2T) human reference genome to systematically map LRC structure across all human chromosomes. Most chromosomes exhibited persistent long-range dependence (LRD), with chromosome-specific scaling intervals spanning kilobase to megabase scales and wavelet Hurst exponents consistently above the uncorrelated expectation of 0.5. Spatially resolved Hurst landscapes revealed that this signal is not uniformly distributed along chromosomes, but instead reflects a heterogeneous genomic mosaic. For the purine-pyrimidine encoding, chromosomes 9, 15, X, and Y showed complex fluctuation profiles not adequately summarized by a single scaling exponent; multifractal detrended fluctuation analysis (MF-DFA) confirmed broad, q-dependent multiscaling consistent with multifractal behavior in these chromosomes, consistent with heterogeneous contributions from satellite-rich and structurally distinct sequence compartments. Surrogate analyses further showed that the observed correlations exceed expectations from base composition alone and are not fully explained by local sequence structure preserved in block-shuffled controls. However, short-memory Autoregressive Moving-Average (ARMA) surrogates reproduced the observed H range in a subset of chromosomes, indicating that the strength of evidence for LRD is chromosome-dependent. Together, these results demonstrate that LRC is a widespread but spatially heterogeneous property of the complete human genome. The T2T assembly reveals that previously unresolved repetitive and satellite-rich regions are strongly associated with chromosome-scale and local scaling variation, providing a framework for future studies of genome evolution, chromatin structure, recombination, structural variation, and genomic instability.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.