P-KNN: Maximizing variant classification evidence through joint calibration of multiple pathogenicity prediction tools
Lin, P.-Y.; Brandes, N.
Show abstract
Clinical guidelines for Mendelian disease diagnosis require that outputs from variant pathogenicity prediction tools be converted into well-calibrated probabilities. However, the existing calibration process is only valid when pre-committing to a specific tool, preventing clinicians from using multiple tools with complementary strengths. To lift this restriction, we introduce Pathogenicity K-Nearest Neighbors (P-KNN), a simple, flexible method that jointly calibrates any set of tools. P-KNN scores each variant by the fraction of pathogenic variants among those with the most similar scores across all underlying tools. We compared P-KNN to single-tool calibration over 13 real tools, including two meta-predictors. Compared to BayesDel, the best-performing tool, P-KNN produced better-calibrated probabilities and stronger evidence strength, with mean log likelihood ratios of 2.56 vs. 2.24 for pathogenic variants, 2.29 vs. 1.56 for benign variants, and 2.28 vs. 2.07 for variants of uncertain significance. We also evaluated P-KNN at four historical time points to assess scalability and robustness, finding that it improved steadily as new tools became available, while remaining well-calibrated. It also correctly integrated computational predictions with experimental measurements, whereas guidelines-based evidence summation systematically overestimated pathogenicity. In summary, P-KNN allows full flexibility to work with any set of tools in a reliable and robust manner. It is consistent with clinical variant classification guidelines, making it well positioned to improve diagnostic yield for rare genetic diseases. P-KNN is available via command line (https://github.com/Brandes-Lab/P-KNN) and precomputed scores (Dataset: https://huggingface.co/datasets/brandeslab/P-KNN, User Interface: https://huggingface.co/spaces/brandeslab/P-KNN-Viewer).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 97%
- A probabilistic graphical model for estimating selection coefficient of nonsynonymous variants from human population sequence data 96%
- ProSolo: Accurate Variant Calling from Single Cell DNA Sequencing Data 95%
Similar papers in this journal
- Genomics 2 Proteins portal: A resource and discovery tool for linking genetic screening outputs to protein sequences and structures 94%
- Merfin: improved variant filtering and polishing via k-mer validation 94%
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 94%
Similar papers in this journal
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 97%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 95%
- Pan-cancer detection of driver genes at the single-patient resolution 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.