Back

Comparative genomics of Tandem Repeat variation in apes

de Lima Adam, C.; Rocha, J. L.; Sudmant, P. H.; Rohlfs, R.

2026-01-21 evolutionary biology
10.64898/2026.01.20.700717 bioRxiv
Show abstract

Tandem repeats (TRs) are highly mutable DNA elements that comprise nearly 8% of the human genome and influence gene regulation, protein coding, and disease. Despite their functional importance, our understanding of TR evolution remains limited, as their repetitive nature has hindered accurate sequencing, annotation, and cross-species comparison. As a result, we lack a population-aware evolutionary framework to quantify TR conservation, divergence, and mutational dynamics across species. Here, using telomere-to-telomere (T2T) reference genomes for seven ape species and population-scale long reads for humans and chimpanzees, we generated a comprehensive comparative catalog of STRs and VNTRs. We identified over 3 million TR loci per ape genome and nearly 2 million homologous loci between humans and chimpanzees. TR diversity and conservation are strongly structured by genomic context, with coding and untranslated regions exhibiting reduced polymorphism and divergence, while intronic and intergenic regions show elevated variability. Heterozygosity varies systematically across species, motif lengths, and functional categories, and mutation rates show strong concordance between indirect and pedigree-based estimates. Using a divergence-diversity ratio framework, we identified TRs under extreme evolutionary regimes that are enriched in genes involved in nervous system development, synaptic function, and cell signaling. Together, these results establish a population and species-resolved framework for studying TR evolution and interpreting TR variation in functional contexts.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.