Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing
Sanchis-Juan, A.; Mostovoy, Y.; Stenton, S. L.; Ganesh, V. S.; Weisburd, B.; Yenkin, A.; Kurtas, N. E.; Zhao, X.; Shin, E.; Boone, P. M.; Su, H.; Lee, A. S.; Yadav, R.; Allan, K.; Argilli, E.; Austin-Tse, C.; Barry, B. J.; Baxter, S.; Beggs, A. H.; Bell, K. M.; Blankenmeister, B.; Bönnemann, C. G.; Brownstein, C. A.; Bujakowska, K. M.; Carbonell, E.; Cooper, S. T.; Covill, L. E.; DiTroia, S.; Donkervoort, S.; Engle, E. C.; Gallacher, L.; Genetti, C. A.; Gleeson, J. G.; Guan, B.; Hall, S.; Hildebrandt, F.; Hufnagel, R. B.; Jurgens, J. A.; Khorgade, A.; Lemire, G.; Liau, E.; Ma, J.; Madden, J.
Show abstract
Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Detecting cryptic clinically-relevant structural variation in exome sequencing data increases diagnostic yield for developmental disorders 96%
- Genome Sequencing and Comprehensive Rare Variant Analysis of 465 Families with Neurodevelopmental Disorders 96%
- Transcriptome-wide outlier approach identifies individuals with minor spliceopathies 96%
Similar papers in this journal
- IGenomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes 98%
- Poison exon annotations improve the yield of clinically relevant variants in genomic diagnostic testing 96%
- Diagnosing missed cases of spinal muscular atrophy in genome, exome, and panel sequencing datasets 96%
Similar papers in this journal
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 97%
- STRchive: a dynamic resource detailing population-level and locus-specific insights at tandem repeat disease loci 96%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 96%
Similar papers in this journal
- Diagnostic Utility of Genome-wide DNA Methylation Analysis in Genetically Unsolved Developmental and Epileptic Encephalopathies and Refinement of a CHD2 Episignature 96%
- Functional annotation of rare structural variation in the human brain 95%
- Exome-wide analysis of congenital kidney anomalies reveals new genes and shared architecture with developmental disorders 95%
Similar papers in this journal
- Comprehensive reanalysis for CNVs in ES data from unsolved rare disease cases results in new diagnoses 95%
- Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort 94%
- Whole genome sequencing delineates regulatory and novel genic variants in childhood cardiomyopathy 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.