Back

Benchmarking of variant pathogenicity prediction methods using a population genetics approach

Gudkov, M.; Thibaut, L.; Monger, S.; Das, D.; Congenital Heart Disease Synergy Study group, ; Winlaw, D. S.; Dunwoodie, S. L.; Giannoulatou, E.

2025-03-17 genomics
10.1101/2025.03.16.643565 bioRxiv
Show abstract

MotivationVariant pathogenicity predictors are essential for identifying new associations between genetic variants and rare diseases. However, despite the availability of numerous predictors, there is no clear consensus on which methods provide the most reliable results. The common practice of training, testing, and benchmarking these predictors using known variant sets from disease or mutagenesis studies raises concerns about ascertainment bias and data circularity. ResultsWe benchmarked commonly used pathogenicity predictors using an orthogonal approach that does not rely on predefined "ground truth" datasets. By leveraging population-level genomic data from gnomAD and the Context-Adjusted Proportion of Singletons (CAPS) metric, we identified CADD and REVEL as the best-performing predictors for distinguishing extremely deleterious variants from moderately deleterious ones. REVEL demonstrated superior calibration. Additionally, we show that CAPS can serve as a meta-analysis tool for interpreting variant annotations and highlight biases in ClinVar-based predictor training. Availability and ImplementationCAPS analysis and benchmarking results are available at https://github.com/mgudVCCRI/PopGenVariantFiltering Contacte.giannoulatou@victorchang.edu.au

Published in Bioinformatics Advances (predicted rank #23) · training set

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.