Back

High rate of mutation and efficient removal by selection of structural variants from natural populations of Caenorhabditis elegans

Saxena, A. S.; Baer, C. F.

2025-03-25 genetics
10.1101/2025.03.22.644739 bioRxiv
Show abstract

The importance of genomic structural variants (SVs) is well-appreciated, but much less is known about their mutational properties than of single nucleotide variants (SNVs) and short indels. The reason is simple: the longer the variant, the less likely it will be covered by a single sequencing read, thus the harder it is to map unambiguously to a unique genomic location. Here we report SV mutation rate estimates from six mutation accumulation (MA) lines from two strains of C. elegans (N2 and PB306) using long-read (PacBio) sequencing. The inferred SV mutation rate is [~]1/10 the SNV rate and [~]1/4 the short indel rate. We identified 40 mutations and 52 false positive (FP) variants by manual inspection of each SV call. Excluding one atypical line (5 mutations, 35 FPs), the signal (mutant) to noise (FP) ratio is approximately 2:1. False negative rates were determined by simulating variants in the reference genome and observing recall. Recall rate ranges from >90% for short indels and declines as SV length increases. Small deletions have nearly the same recall rate as small insertions ([~]100bp), but deletions have higher recall rates than insertions as size increases. The reported SV mutation rate is likely a lower bound. A quarter of identified SV mutations occur in SV hotspots that harbor pre-existing low complexity repeat variation. Comparison of the spectrum of spontaneous SVs to wild isolates implies that natural selection is not only efficient at removing SVs in exons but also removes roughly half of SVs in intergenic regions.

Published in Genome Biology and Evolution (predicted rank #5) · training set

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.