Back

Direct identification of de novo mobile element insertions from single molecule sequencing of human sperm

Li, S.; Gozashti, L.; Connelly, C.; Goubert, C.; Aston, K.; Gleeson, J. G.; Quinlan, A.; Yang, X.; Sudmant, P. H.

2026-08-26 genetics
10.1101/2025.10.25.684559 bioRxiv
Show abstract

Mobile element insertions (MEIs) are a significant source of human genetic variation, yet the rates and properties of de novo MEIs are poorly characterized due to technical limitations in sequencing technology. Here, we directly sequenced individual gametes from sperm samples of 19 donors (aged 27-62) using highly accurate PacBio long-read sequencing to identify de novo retrotransposition events without familial inference. We developed a "self-alignment" strategy using personalized genome assemblies that enables high-precision, single-read detection of de novo MEIs. Using this method, we identified 43 de novo Alu insertions, revealing >9-fold variation in Alu retrotransposition rates between individuals (ranging from 0 to 0.148 insertions/gamete). We found a significant increase in Alu activity with paternal age, yielding a 4.67% increase in insertions per gamete per year of additional paternal age, representing a direct observation of age-associated increases in structural variant (SV) mutation rates. De novo Alu insertions predominantly represent evolutionarily young AluYa5 and AluYb8 subfamilies and bear characteristic molecular signatures of target-primed reverse transcription (TPRT). Our population-averaged rate of 4.52 insertions per 100 gametes aligns well with previous population genetic estimates, validating both direct observation and population approaches for estimating de novo MEI rates. These results establish direct gamete sequencing as a powerful method for characterizing germline mutation processes and reveal age as a significant determinant of de novo retrotransposition in the male germline.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.