Multimodal framework to resolve variants of uncertain significance in TSC2
Biar, C. G.; Pfeifer, C.; Carvill, G. L.; Calhoun, J. D.
Show abstract
Efforts to resolve the functional impact of variants of uncertain significance (VUS) have lagged behind the identification of new VUS; as such, there is a critical need for scalable VUS resolution technologies. Computational variant effect predictors (VEPs), once trained, can predict pathogenicity for all missense variants in a gene, set of genes, or the exome. Existing tools have employed information on known pathogenic and benign variants throughout the genome to predict pathogenicity of VUS. We hypothesize that taking a gene-specific approach will improve pathogenicity prediction over globally-trained VEPs. We tested this hypothesis using the gene TSC2, whose loss of function results in tuberous sclerosis, a multisystem mTORopathy affecting about 1 in 6,000 individuals born in the United States. TSC2 has been identified as a high-priority target for VUS resolution, with (1) well-characterized molecular and patient phenotypes associated with loss-of-function variants, and (2) more than 2,700 VUS already documented in ClinVar. We developed Tuberous sclerosis classifier to Resolve variants of Uncertain Significance in TSC2 (TRUST), a machine learning model to predict pathogenicity of TSC2 missense VUS. To test whether these predictions are accurate, we further introduce curated loci prime editing (cliPE) as an accessible strategy for performing scalable multiplexed assays of variant effect (MAVEs). Using cliPE, we tested the effects of more than 200 TSC2 variants, including 106 VUS. It is highly likely this functional data alone would be sufficient to reclassify 92 VUS with most being reclassified as likely benign. We found that TRUSTs classifications were correlated with the functional data, providing additional validation for the in silico predictions. We provide our pathogenicity predictions and MAVE data to aid with VUS resolution. In the near future, we plan to host these data on a public website and deposit into relevant databases such as MAVEdb as a community resource. Ultimately, this study provides a framework to complete variant effect maps of TSC1 and TSC2 and adapt this approach to other mTORopathy genes.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-throughput splicing assays identify missense and silent splice-disruptive POU1F1 variants underlying pituitary hormone deficiency 95%
- Dystonia-specific mutations in THAP1 alter transcription of genes associated with neurodevelopment and myelin 94%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 94%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 94%
- Identification and validation of novel candidate risk genes in endocytic vesicular trafficking associated with esophageal atresia and tracheoesophageal fistulas 93%
- Functional characterization of pathogenic SATB2 missense variants identifies distinct effects on chromatin binding and transcriptional activity 93%
Similar papers in this journal
- ParSE-seq: A Calibrated Multiplexed Assay to Facilitate the Clinical Classification of Putative Splice-altering Variants 94%
- Diagnostic Utility of Genome-wide DNA Methylation Analysis in Genetically Unsolved Developmental and Epileptic Encephalopathies and Refinement of a CHD2 Episignature 94%
- Multi-model functionalization of disease-associated PTEN missense mutations identifies multiple molecular mechanisms underlying protein dysfunction 94%
Similar papers in this journal
Similar papers in this journal
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 95%
- Systematic analysis of genetic and phenotypic characteristics reveals antisense oligonucleotide therapy potential for one-third of neurodevelopmental disorders 94%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.