Back

Single-molecule variation in telomeric sequence and structure across humans

Dubocanin, D.; Vollger, M. R.; Neph, S. J.; Del Rio Pisula, M.; Lucas, J. K.; Sedeno-Cortes, A. E.; Mallory, B. J.; Real, T. D.; Human Pangenome Reference Consortium, ; Barthel, F. P.; Altemose, N.; Stergachis, A. B.

2026-05-05 genomics
10.64898/2026.05.01.722200 bioRxiv
Show abstract

The repetitive architectures of telomeric and subtelomeric regions have obscured studies of their genetic variation and chromatin organization across the human population. Here, we integrate near-complete diploid genome assemblies from 212 individuals with matched long-read sequencing data to construct an atlas of 316,146 telomere-spanning molecules across 12,080 chromosome-end-resolved telomere arrays. This atlas reveals that nearly every chromosome end harbors a structured and unique pattern of telomere variant repeats (TVR), or TVR code, with subtelomere-proximal TVR codes being heritable, somatically stable, and influenced by subtelomeric TAR1 regulatory elements. Despite ongoing cycles of telomere shortening and elongation in the germline, proximal TVR codes are maintained across the human population. These TVR codes expose rare telomerase-independent events that lengthen telomeres in the germline, including interchromosomal telomere exchange and recurrent internal duplications within telomere arrays. Furthermore, single-molecule chromatin fiber sequencing across 26,972 molecules spanning the telomere-subtelomere boundary confirms that TVR-rich regions adopt telomeric chromatin but introduce discrete discontinuities into otherwise compact telomeric chromatin fibers. Together, our results link chromosome-end sequence variation to telomere cap formation and telomerase-independent telomere extension mechanisms in the human germline.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.