Expectations and blind spots for structural variation detection from short-read alignment and long-read assembly
Zhao, X.; Collins, R. L.; Lee, W.-P.; Weber, A. M.; Jun, Y.; Zhu, Q.; Weisburd, B.; Huang, Y.; Audano, P. A.; Wang, H.; Walker, M.; Lowther, C.; Fu, J.; Gerstein, M. B.; Devine, S. E.; Marschall, T.; Korbel, J. O.; Eichler, E. E.; Chaisson, M. J. P.; Lee, C.; Mills, R. E.; Brand, H.; Talkowski, M. E.
Show abstract
Virtually all genome sequencing efforts in national biobanks, complex and Mendelian disease programs, and emerging clinical diagnostic approaches utilize short-reads (srWGS), which present constraints for genome-wide discovery of structural variants (SVs). Alternative long-read single molecule technologies (lrWGS) offer significant advantages for genome assembly and SV detection, while these technologies are currently cost prohibitive for large-scale disease studies and clinical diagnostics (∼5-12X higher cost than comparable coverage srWGS). Moreover, only dozens of such genomes are currently publicly accessible by comparison to millions of srWGS genomes that have been commissioned for international initiatives. Given this ubiquitous reliance on srWGS in human genetics and genomics, we sought to characterize and quantify the properties of SVs accessible to both srWGS and lrWGS to establish benchmarks and expectations in ongoing medical and population genetic studies, and to project the added value of SVs uniquely accessible to each technology. In analyses of three trios with matched srWGS and lrWGS from the Human Genome Structural Variation Consortium (HGSVC), srWGS captured ∼11,000 SVs per genome using reference-based algorithms, while haplotype-resolved assembly from lrWGS identified ∼25,000 SVs per genome. Detection power and precision for SV discovery varied dramatically by genomic context and variant class: 9.7% of the current GRCh38 reference is defined by segmental duplications (SD) and simple repeats (SR), yet 91.4% of deletions that were specifically discovered by lrWGS localized to these regions. Across the remaining 90.3% of the human reference, we observed extremely high concordance (93.8%) for deletions discovered by srWGS and lrWGS after error correction using the raw lrWGS reads. Conversely, lrWGS was superior for detection of insertions across all genomic contexts. Given that the non-SD/SR sequences span 90.3% of the GRCh38 reference, and encompass 95.9% of coding exons in currently annotated disease associated genes, improved sensitivity from lrWGS to discover novel and interpretable pathogenic deletions not already accessible to srWGS is likely to be incremental. However, these analyses highlight the added value of assembly-based lrWGS to create new catalogues of functional insertions and transposable elements, as well as disease associated repeat expansions in genomic regions previously recalcitrant to routine assessment.Competing Interest StatementThe authors have declared no competing interest.View Full Text
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 96%
- Long read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits 96%
- Inferring compound heterozygosity from large-scale exome sequencing data 95%
Similar papers in this journal
Similar papers in this journal
- Comprehensive analysis of structural variants in breast cancer genomes using single molecule sequencing 95%
- Characterising tandem repeat complexities across long-read sequencing platforms with TREAT and otter 95%
- A comprehensive catalog of 3D genome organization in diverse human genomes facilitates understanding of the impact of structural variation on chromatin structure 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.