Systematic identification of disease-causing promoter and untranslated region variants in 8,040 undiagnosed individuals with rare disease
Martin Geary, A. C.; Blakes, A. J.; Dawes, R.; Findlay, S. D.; Lord, J. C.; Walker, S.; Talbot-Martin, J.; Wieder, N.; D'Souza, E. N.; Fernandes, M.; Hilton, S.; Lahiri, N.; Campbell, C.; Jenkinson, S.; De Goede, C. G.; Anderson, E. R.; Burge, C. B.; Sanders, S. J.; Ellingford, J.; Baralle, D.; Banka, S.; Whiffin, N.
Show abstract
BackgroundBoth promoters and untranslated regions (UTRs) have critical regulatory roles, yet variants in these regions are largely excluded from clinical genetic testing due to difficulty in interpreting pathogenicity. The extent to which these regions may harbour diagnoses for individuals with rare disease is currently unknown. MethodsWe present a framework for the identification and annotation of potentially deleterious proximal promoter and UTR variants in known dominant disease genes. We use this framework to annotate de novo variants (DNVs) in 8,040 undiagnosed individuals in the Genomics England 100,000 genomes project, which were subject to strict region-based filtering, clinical review, and validation studies where possible. In addition, we performed region and variant annotation-based burden testing in 7,862 unrelated probands against matched unaffected controls. ResultsWe prioritised eleven DNVs and identified an additional variant overlapping one of the eleven. Ten of these twelve variants (82%) are in genes that are a strong match to the individuals phenotype and six had not previously been identified. Through burden testing, we did not observe a significant enrichment of potentially deleterious promoter and/or UTR variants in individuals with rare disease collectively across any of our region or variant annotations. ConclusionsOverall, we demonstrate the value of screening promoters and UTRs to uncover additional diagnoses for previously undiagnosed individuals with rare disease and provide a framework for doing so without dramatically increasing interpretation burden.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 95%
- MRSD: a novel quantitative approach for assessing suitability of RNA-seq in the clinical investigation of mis-splicing in Mendelian disease 95%
- Non-coding variants upstream of MEF2C cause severe developmental disorder through three distinct loss-of-function mechanisms 94%
Similar papers in this journal
- Systematic analysis of genetic and phenotypic characteristics reveals antisense oligonucleotide therapy potential for one-third of neurodevelopmental disorders 96%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 95%
- Clustering of predicted loss-of-function variants in genes linked with monogenic disease can explain incomplete penetrance 94%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 95%
- Evaluation of imputation performance of multiple reference panels in a Pakistani population 94%
- BinomiRare: A carriers-only test for association of rare genetic variants with a binary outcome for mixed models and any case-control proportion 93%
Similar papers in this journal
Similar papers in this journal
- Deciphering novel TCF4-driven mechanisms underlying a common triplet repeat expansion-mediated disease 95%
- Missense variants causing Wiedemann-Steiner syndrome preferentially occur in the KMT2A-CXXC domain and are accurately classified using AlphaFold2 94%
- Genome mining yields new disease-associated ROMK variants with distinct defects 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.