The Impact of Clinical and Molecular Variant Properties on Calibration and Performance of Variant Effect Prediction Tools
Isakov, O.; Ben-Shachar, S.
Show abstract
BackgroundVariant Effect Prediction (VEP) tools are essential for determining the potential pathogenicity of genetic variants, aiding clinical diagnostics and genetic counseling. However, their performance can vary depending on molecular and clinical contexts, complicating variant classification. AimThis study aims to assess the performance variability of commonly used VEP tools under different conditions. Additionally, the study aims to recalibrate score thresholds to better reflect evidence of pathogenicity. MethodsClinVar variants classified as pathogenic (P), likely pathogenic (LP), benign (B), or likely benign (LB) were analyzed using 25 VEP tools. Tools were evaluated based on discriminatory performance. Data were stratified by variant creation date, allele frequency, conservation level, mode of inheritance (MOI), and disease category. For each subset, Bayesian methods were employed to recalibrate score thresholds corresponding to the levels of evidence defined by the American College of Medical Genetics (ACMG). ResultsThe performance of VEP tools varied significantly across different subsets. Variants created after 2020 showed a mild yet significant decrease in the performance of certain VEP tools, particularly those trained on earlier datasets. VEP tools exhibited reduced accuracy for variants with higher allele frequencies, particularly those exceeding a frequency of 10-4, suggesting a need for recalibration when assessing more common variants. Tools demonstrated lower discriminatory performance for variants located in regions with high conservation, mostly due to a decrease in specificity. Variants affecting autosomal recessive (AR) and X-linked (XL) genes were more accurately classified compared to those affecting autosomal dominant (AD) genes by most tools. Differences in MOI and conservation levels within certain disease categories were shown to correlate with overall performance. Recalibration of prediction scores resulted in lower score thresholds in low conservation regions compared to high and uncovered subsets in which higher levels of evidence may be achieved. ConclusionVEP tools exhibit context-dependent performance variability, necessitating score recalibration for accurate classification.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome Alert!: a standardized procedure for genomic variant reinterpretation and automated genotype-phenotype reassessment in clinical routine 95%
- Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms 95%
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 94%
Similar papers in this journal
- ConsensuSV-ONT - a modern method for accurate structural variant calling 93%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 93%
- GeneToCN: An Alignment-Free Method for Gene Copy Number Estimation Directly from Next-Generation Sequencing Reads 93%
Similar papers in this journal
- Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation framework 94%
- VarSight: Prioritizing Clinically Reported Variants with Binary Classification Algorithms 93%
- ATAV: a comprehensive platform for population-scale genomic analyses 92%
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 95%
- OncoGEMINI: Software for Investigating Tumor Variants From Multiple Biopsies With Integrated Cancer Annotations 93%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.