Back

Cross-species discovery of large structural variants reveals distinct evolutionary dynamics in primates

Zhou, F.; Han, J.; Zhang, S.; Lin, J.; Eichler, E.; Mao, Y.

2026-08-28 genomics
10.64898/2026.08.28.747789 bioRxiv
Show abstract

Large-scale structural variants (SVs) are major drivers of genome evolution and disease susceptibility, shaping lineage-specific gene content, genome architecture, and phenotypic diversity. Yet, their systematic detection across divergent species has been hindered by complex rearrangements and reference bias, limiting our understanding of their evolutionary impact. Here we present LGvar, a divergence-aware tool for identifying SVs in population-scale and cross-species genome assemblies, enabling up to a 7-fold increase in the detection of large deletions, insertions, inversions and complex rearrangements ([&ge;]10 kbp). Application of LGvar uncovers 112 previously undetected large SVs overlooked by conventional methods and quantifies the effect of evolutionary divergence on SV discovery across primate euchromatic genomes, revealing an approximately 30% overall decline in detection sensitivity with increasing phylogenetic distance. Mapping SVs from comparable syntenic regions onto the primate phylogeny shows that large deletions ([&ge;]10 kbp) accumulate approximately two-fold faster than insertions and eight-fold faster than inversions, in contrast to the dynamics of smaller SVs (<10 kbp). Integrating structurally divergent regions with the synteny region analysis reveals over 606 previously unreported gene gains and losses in great apes and refined the totals to 134 human-lineage-specific and 51 great-ape-lineage-specific protein coding genes. These results suggest that large SVs follow distinct evolutionary trajectories, with potential selective constraints shaping genome structure. Our study establishes a generalizable strategy for cross-species SV discovery and provides a high-resolution view of how large SVs contribute to genome evolution in primates.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.