SNPred outperforms other ensemble-based SNV pathogenicity predictors and elucidates the challenges of using ClinVar for evaluation of variant classification quality.
Molotkov, I.; Koboldt, D.; Artomov, M.
Show abstract
BackgroundCurrent single nucleotide variants (SNVs) pathogenicity prediction tools assess various properties of genetic variants and provide a likelihood of causing a disease. This information aids in variant prioritization - the process of narrowing down the list of potential pathogenic variants, and, therefore, facilitating diagnostics. Assessing the effectiveness of SNV pathogenicity tools using ClinVar data is a widely adopted practice. Our findings demonstrate that this conventional method tends to overstate performance estimates. MethodsWe introduce SNPred, an ensemble model specifically designed for predicting the pathogenicity of nonsynonymous single nucleotide variants (nsSNVs). To evaluate its performance, we conducted assessments using six distinct validation datasets derived from ClinVar and BRCA1 Saturation Genome Editing (SGE) data. ResultsAcross all validation scenarios, SNPred consistently outperformed other state-of-the-art tools, particularly in the case of rare and cancer-related variants, as well as variants that are classified with low confidence by most in silico tools. To ensure convenience, we provide precalculated scores for all possible nsSNVs. We proved that the exceptionally high accuracy scores of the best models achieved for ClinVar variants are only attainable if the models learn to replicate misclassifications found in ClinVar. Additionally, we conducted a comparison of predictor performance on two distinct sets of BRCA1 variants that did not overlap: one sourced from ClinVar and the other from the SGE study. Across all in silico predictors, we observed a significant trend where ClinVar variants were classified with notably higher accuracy. ConclusionsWe provide a powerful variant pathogenicity predictor that enhances the quality of clinical variant interpretation and highlights important challenges of using ClinVar for SNV pathogenicity predictors evaluation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome Alert!: a standardized procedure for genomic variant reinterpretation and automated genotype-phenotype reassessment in clinical routine 96%
- Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms 95%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 95%
Similar papers in this journal
- Characteristics predicting reduced penetrance variants in the high-risk cancer predisposition gene TP53 94%
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 93%
- Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders 93%
Similar papers in this journal
- Evidence-based calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for clinical use of PP3/BP4 criteria 97%
- Evidence-based recommendations for gene-specific ACMG/AMP variant classification from the ClinGen ENIGMA BRCA1 and BRCA2 Variant Curation Expert Panel 95%
- Availability of benign missense variant “truthsets” for validation of functional assays: current status and a novel systematic approach 95%
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 97%
- varCADD: large sets of standing genetic variation enable genome-wide pathogenicity prediction 94%
- OncoGEMINI: Software for Investigating Tumor Variants From Multiple Biopsies With Integrated Cancer Annotations 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.