Random forest model improves annotation and discovery of variants of uncertain significance in Alzheimer's and other neurological disorders
Jonson, C. P.; Makarious, M. B.; Koretsky, M. J.; Vitale, D.; Leonard, H. L.; Rabkina, L.; Lario Lago, A.; Eslami-Amirabadi, M.; Ramos, E. M.; Narayan, P.; Yokoyama, J. S.; Singleton, A. B.; Blauwendraat, C.; Cookson, M. R.; Nalls, M.
Show abstract
Variants of uncertain significance (VUS) are a bottleneck for genetic discovery and complicate clinical decision-making in Alzheimers disease and related neurological disorders (ADRD). We developed MoVUS: Model for Variants of Unknown Significance, a random-forest approach that integrates functional predictors to classify missense VUS. MoVUS leverages a balanced random forest model trained on dbNSFP v5.1a with high-confidence ClinVar and HGMD labels, using harmonized functional prediction rankscores. MoVUS produced confident, explainable calls, with [~]98% accuracy (AUC [~]0.998), prioritizing potentially pathogenic candidates and down-ranked likely benign variants on independent validation sets of ClinVar-only and HGMD-only variants. In our discovery analyses on ADRD-implicated variants in dbNSFP and from independent collaborator cohorts, we achieved high-confidence classifications on a majority of the unknown variants (average of 55% of discovery variants). We also had access to medical records and family trees for some variants, further validating our findings. Across held-out and external datasets, MoVUS reports high accuracy alongside confidence scores and helps prioritize actionable candidates, and reduces bias by considering multiple scores for each variant. To facilitate use, we developed a web app for users to browse across 100+ ADRD genes. MoVUS provides transparent, reproducible triage for research follow-up by pairing consensus predictors with SHAP-based visualizations and explanations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 96%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 94%
- Poison exon annotations improve the yield of clinically relevant variants in genomic diagnostic testing 94%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 94%
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 94%
- A Specialized Reference Panel with Structural Variants Integration for Improving Genotype Imputation in Alzheimer's Disease and Related Dementias (ADRD) 94%
Similar papers in this journal
- Non-Coding and Loss-of-Function Coding Variants in TET2 are Associated with Multiple Neurodegenerative Diseases 94%
- Evidence-based calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for clinical use of PP3/BP4 criteria 94%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 94%
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 95%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 95%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 94%
Similar papers in this journal
- A fast and robust strategy to remove variant level artifacts in Alzheimer’s Disease Sequencing Project data 96%
- A 3’ UTR Deletion Is a Leading Candidate Causal Variant at the TMEM106B Locus Reducing Risk for FTLD-TDP 95%
- Association of genetic variation at the GJA5/ACP6 locus with motor progression in Parkinson’s 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.