Gene- and domain-aware calibration increases the clinical utility of variant effect predictors
Chen, Y.; Fayer, S.; Jain, S.; Benazouz, M.; Sverchkov, Y.; Stone, J.; Sharma, H.; Bergquist, T.; Stewart, R.; Mooney, S. D.; Craven, M.; Radivojac, P.; Starita, L. M.; Fowler, D. M.; Pejaver, V.
Show abstract
Approximately 90% of missense variants in ClinVar are variants of uncertain significance, limiting clinical utility of genetic testing. Variant effect predictors (VEPs) generate scores for any missense variant, providing massive potential to empower classification. Realizing this potential requires calibration to translate VEP scores into evidence. However, current genome-wide calibration masks predictor heterogeneity across genes, causing evidence misassignment. We developed an automated and flexible calibration framework performing gene-specific calibration when control variants are sufficient. For genes with limited control variants, we aggregated protein domains with similar VEP score distributions to enable robust calibration. Applying this framework to three VEPs across 2,769 genes, gene-specific and domain-based calibration increased variants with assigned evidence and improved evidence accuracy versus genome-wide calibration. These calibrations are available through PredictMD (igvf.mavedb.org), providing clinicians with calibrated computational evidence. Together these calibration strategies substantially increase the clinical utility of VEPs for variant classification.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 97%
- A probabilistic graphical model for estimating selection coefficient of nonsynonymous variants from human population sequence data 96%
- ProSolo: Accurate Variant Calling from Single Cell DNA Sequencing Data 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.