The impact of systematized generation, evaluation, and incorporation of machine learning algorithms for clinical variant classification
Fresard, L.; Facio, F. M.; Chen, E.; Colavin, A.; Johnson, B.; Araya, C.; Manders, T.; Wahl, A.; Metz, H.; Nicoludis, J. M.; Ouyang, K.; Padigepati, S.; Kobayashi, Y.; Reuter, J.; Nykamp, K.
Show abstract
Variants of uncertain significance (VUS) pose a significant challenge for those undergoing genetic testing, leading to prolonged uncertainty and inappropriate medical care. VUS rate reduction is critical to fully realize the utility of genetic testing for all populations. With the growth of large-scale biological data sources and modern Machine Learning (ML) techniques, predictive modeling has enormous potential for VUS reduction. For this purpose, we developed the Invitae Evidence Modeling Platform (EMP), with key features designed to maximize the utility and confidence of predictive algorithms for variant classification. First, input data for a new model is curated to correspond to a single major evidence category within a variant classification framework. Second, gene-specific training and/or validation is performed for each model type. Third, accuracy thresholds are set to filter out gene-specific models that do not meet stringent accuracy metrics. Finally, prediction scores for variant pathogenicity are calibrated to ensure internally consistent evidence weighting within the classification framework. The EMP has accelerated the development of ML algorithms and greatly expanded the amount of evidence available for variant classification. EMP evidence has been applied to more than 800,000 variants across 1 million individuals, 42% of which would have been VUS without this evidence. Importantly, definitive classifications (P, LP, LB, B) made with EMP evidence have high prospective concordance (>99%) with ClinVar submissions. Finally, we demonstrate that further use and development of EMP evidence for variant classification has the potential to reduce the VUS disparity across race/ethnicity/ancestry (REA) groups.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evidence-based calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for clinical use of PP3/BP4 criteria 97%
- Evidence-based recommendations for gene-specific ACMG/AMP variant classification from the ClinGen ENIGMA BRCA1 and BRCA2 Variant Curation Expert Panel 97%
- Advanced variant classification framework reduces the false positive rate of predicted loss of function (pLoF) variants in population sequencing data 97%
Similar papers in this journal
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 97%
- Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms 97%
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 96%
Similar papers in this journal
- When splicing is not all or none: Implications for variant classification 94%
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 94%
- Characteristics predicting reduced penetrance variants in the high-risk cancer predisposition gene TP53 94%
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 96%
- From Text to Translation: Using Language Models to Prioritize Variants for Clinical Review 95%
- Recommendations for clinical interpretation of variants found in non-coding regions of the genome 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.