Back

Real-world evaluation of deep learning algorithms to classify functional pathogenic germline variants

Chow, R. D.; Parikh, R. B.; Nathanson, K. L.

2024-04-07 genetic and genomic medicine
10.1101/2024.04.05.24305402 medRxiv
Show abstract

Deep learning models for variant pathogenicity prediction can recapitulate expert-curated annotations, but their performance remains unexplored on actual disease phenotypes in a real-world setting. Here, we apply three state-of-the-art pathogenicity prediction models to classify hereditary breast cancer gene variants in the UK Biobank. Predicted pathogenic variants in BRCA1, BRCA2 and PALB2, but not ATM and CHEK2, were associated with increased breast cancer risk. We explored gene-specific score thresholds for variant pathogenicity, finding that they could improve model performance. However, when specifically tasked with classifying variants of uncertain significance, the deep learning models were generally of limited clinical utility.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.