Long-read sequencing resolves the clinically relevant CYP21A2 locus, supporting a new clinical test for Congenital Adrenal Hyperplasia
Monlong, J.; Chen, X.; Barseghyan, H.; Rowell, W. J.; Negi, S.; Nokoff, N.; Mohnach, L.; Hirsch, J.; Finlayson, C.; Keegan, C. E.; Almalvez, M.; Berger, S. I.; DeDios, I.; McNulty, B.; Robertson, A.; Miga, K.; Speiser, P. W.; Paten, B.; Vilain, E.; Delot, E. C.
Show abstract
Congenital Adrenal Hyperplasia (CAH), one of the most common inherited disorders, is caused by defects in adrenal steroidogenesis. It is potentially lethal if untreated and is associated with multiple comorbidities, including fertility issues, obesity, insulin resistance, and dyslipidemia. CAH can result from variants in multiple genes, but the most frequent cause is deletions and conversions in the segmentally duplicated RCCX module, which contains the CYP21A2 gene and a pseudogene. The molecular genetic test to identify pathogenic alleles is cumbersome, incomplete, and available from a limited number of laboratories. It requires testing parents for accurate interpretation, leading to healthcare inequity. Less severe forms are frequently misdiagnosed, and phenotype/genotype correlations incompletely understood. We explored whether emerging technologies could be leveraged to identify all pathogenic alleles of CAH, including phasing in proband-only cases. We targeted long-read sequencing outputs that would be practical in a clinical laboratory setting. Both HiFi-based and nanopore-based whole-genome long-read sequencing datasets could be mined to accurately identify pathogenic single-nucleotide variants, full gene deletions, fusions creating non-functional hybrids between the gene and pseudogene ("30-kb deletion"), as well as count the number of RCCX modules and phase the resulting multimodular haplotypes. On the Hi-Fi data set of 6 samples, the PacBio Paraphase tool was able to distinguish nine different mono-, bi-, and tri-modular haplotypes, as well as the 30-kb and whole gene deletions. To do the same on the ONT-Nanopore dataset, we designed a tool, Parakit, which creates an enriched local pangenome to represent known haplotype assemblies and map ClinVar pathogenic variants and fusions onto them. With few labels in the region, optical genome mapping was not able to reliably resolve module counts or fusions, although designing a tool to mine the dataset specifically for this region may allow doing so in the future. Both sequencing techniques yielded congruent results, matching clinically identified variants, and offered additional information above the clinical test, including phasing, count of RCCX modules, and status of the other module genes, all of which may be of clinical relevance. Thus long-read sequencing could be used to identify variants causing multiple forms of CAH in a single test.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HiFi long-read genomes for difficult-to-detect clinically relevant variants 94%
- The impact of 22q11.2 copy number variants on human traits in the general population 94%
- Identification of actionable genetic variants in 4,198 Scottish volunteers from the Viking Genes research cohort and implementation of return of results 94%
Similar papers in this journal
- Assessment of the variant prioritisation strategy for genomic newborn screening in the Generation Study 95%
- A gene pathogenicity tool 'GenePy' identifies missed biallelic diagnoses in the 100,000 Genomes Project 95%
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 94%
Similar papers in this journal
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 94%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 94%
- Recommendations for clinical interpretation of variants found in non-coding regions of the genome 93%
Similar papers in this journal
- Structural variant calling and clinical interpretation in 6224 unsolved rare disease exomes 93%
- Re-evaluation and Re-analysis of 152 research exomes five years after the initial report reveals clinically relevant changes in 20% 93%
- Limitations in next-generation sequencing-based genotyping of breast cancer polygenic risk score loci 93%
Similar papers in this journal
- Phasing of de novo mutations using a scaled-up multiple amplicon long-read sequencing approach 95%
- GeneBreaker: Variant simulation to improve the diagnosis of Mendelian rare genetic diseases 94%
- Matching whole genomes to rare genetic disorders: Identification of potential causative variants using phenotype-weighted knowledge in the CAGI SickKids5 clinical genomes challenge 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.