Cancer genomic profiling predicts pathogenicity of BRCA1 and BRCA2 variants
Kondrashova, O.; Johnston, R. L.; Parsons, M. T.; Davidson, A. L.; Canson, D. M.; Tran, K. A.; Cline, M. S.; Waddell, N.; Sivakumar, S.; Sokol, E. S.; Jin, D. X.; Pavlick, D. C.; Decker, B.; Frampton, G. M.; Spurdle, A. B.; Parsons, M. T.; Spurdle, A. B.
Show abstract
Accurate classification of BRCA1 and BRCA2 variants is essential for cancer risk assessment and therapy selection, yet over one-third remain variants of uncertain significance (VUS). Here, using 120,660 real-world cancer genomic profiles with BRCA1 or BRCA2 variants from a >800,000-sample cohort, we develop machine learning models that predict pathogenicity using clinical and tumor-derived features, including a pan-cancer homologous recombination deficiency signature, co-mutated genes, zygosity, and cancer type. Trained on classified variants from ClinVar, our models achieved near-perfect performance, with validation ROC-AUC of 1.000 for BRCA1 and 0.989 for BRCA2 variants with [≥]5 observations, translating to strong benign or pathogenic evidence for VCEP classification. Applying these models to 1,073 BRCA1 and 1,639 BRCA2 VUS, we strengthened or enabled classification of 39.48% BRCA1 and 50.52% BRCA2 assessable variants. This approach transforms underutilized tumor profiling data into evidence that can be directly integrated into variant classification, providing a scalable framework for other tumor profiling datasets and cancer genes associated with defined tumor genomic features.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The DiffInvex evolutionary model for conditional somatic selection identifies chemotherapy resistance genes in 10,000 cancer genomes 97%
- Cross-dataset pan-cancer detection: Correlating cell-free DNA fragment coverage with open chromatin sites across cell types 96%
- Integrative ensemble modelling of cetuximab sensitivity in colorectal cancer PDXs 96%
Similar papers in this journal
Similar papers in this journal
- Discovering Monogenic Patients with a Confirmed Molecular Diagnosis in Millions of Clinical Notes with MonoMiner 94%
- Accurate assignment of disease liability to genetic variants using only population data 93%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.