Benchmark for simple and complex genome inversions
Cheng, S.; Sedlazeck, F. J.
Show abstract
BackgroundInversions represent a consequential yet under-characterized form of structural variation, with roles in genomic disorders, evolution, and genome instability. However, their detection remains technically challenging, particularly within repeat-dense regions and for complex multi-breakpoint events. A lack of dedicated, high-quality benchmarks has hindered algorithmic improvement, performance comparison, and robust biological interpretation. ResultsHere, we present a comprehensive, multi-genome benchmark for simple and complex inversions derived from Strand-seq and phased long-read based assemblies across five reference samples, with breakpoint refinement using haplotype-resolved long-read assemblies. This tiered resource spans a broad spectrum of inversion classes, sizes, and zygosity states, and captures challenging genomic contexts including segmental duplications, inverted repeats, and composite rearrangements. We used this benchmark to systematically assess leading structural variant callers and alignment strategies across short-read, PacBio HiFi, and Oxford Nanopore data. Performance varied substantially by inversion class and genomic context: simple inversions were recovered with high sensitivity at sufficient coverage, whereas complex and heterozygous events remained difficult. Sniffles2 and Severus achieved the strongest recall for complex inversions, despite increased false-positive rates. We additionally benchmarked two commonly used long-read alignment pipelines (Minimap2 and VACmap), demonstrating that the mapper choice has a substantial impact on inversion detection in repetitive regions. ConclusionTogether, this work provides the first unified, high-resolution inversion benchmark and reveals clear strengths and limitations of current methods across platforms. Our resource establishes a foundation for principled tool development, evaluation, and tuning, enabling the community to more accurately resolve inversion variation and its biological and clinical consequences.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Gaps and complex structurally variant loci in phased genome assemblies 96%
- Highly accurate assembly polishing with DeepPolisher 95%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
Similar papers in this journal
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 94%
- Complex structural variant visualization with SVTopo 93%
- Towards a better understanding of the low recall of insertion variants with short-read based variant callers 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.