Integrating enriched case data from national laboratory testing with population-based case-control analyses: a novel statistical likelihood-ratio methodology for PS4 applied to 325,345 breast cancer cases and 671,006 controls
Allen, S.; Rowlands, C. F.; Garrett, A.; Couch, F.; Richardson, M. E.; Pesaran, T.; Pethick, J.; Lavelle, K.; McRonald, F.; Vernon, S.; Torr, B.; Loong, L.; Aungraheeta, R.; Durkie, M.; Burghel, G. J.; Callaway, A.; Robinson, R.; Field, J.; Frugtniet, B.; Palmer-Smith, S.; Grant, J.; Pagan, J.; McDevitt, T.; Snape, K.; Hanson, H.; McVeigh, T.; Loveday, C.; Jones, M.; Hardy, S.; Turnbull, C.; CanVIG-UK,
Show abstract
Background: For many evidence criteria within v3.0 of the ACMG/AMP guidelines, methodologies have been developed to empower their use outside the stipulated evidence strengths. However, no such methodology has been established for case-control data (PS4). With the release of large-scale unselected case-control datasets and expansion of nationally-collected laboratory datasets enriched for pathogenic variant carriers, there is potential to combine datasets across ascertainment contexts in a more quantitative manner using novel likelihood ratio tools. Methods: Using our published PS4-LR-Calculator, we calculated a combined log likelihood ratio (PS4-LLR) across five datasets (three unselected, and two enriched), and estimated enrichment of pathogenic variants in clinically-ascertained laboratory data using truncating variant prevalence. Results: Data were combined for 10,817 missense variants from 325,345 female breast cancer patients and 671,006 controls of Western European ancestry for five breast cancer susceptibility genes (BRCA1, BRCA2, PALB2, ATM, CHEK2). A combined LLR was produced for 4,690 missense variants; 927 variants received evidence towards pathogenicity (LLR[≥]1), and 3,242 received evidence towards benignity (LLR[≤]-1). Conclusion: This flexible, variant-level methodology combines nationally-collected 'enriched' datasets with unselected case-control cohorts, expanding the available information for case-control analysis, boosting power, enabling exploration of atypical penetrance and empowering variant classification.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification of Variants of Reduced Penetrance in High Penetrance Cancer Susceptibility Genes: Framework for Genetics Clinicians and Clinical Scientists by CanVIG-UK (Cancer Variant Interpretation Group-UK) 95%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 94%
- Evaluation of Bayesian Classification Framework on the Variant Classification of Hereditary Cancer Predisposition Genes 94%
Similar papers in this journal
- Availability of benign missense variant “truthsets” for validation of functional assays: current status and a novel systematic approach 97%
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 96%
- Evidence-based recommendations for gene-specific ACMG/AMP variant classification from the ClinGen ENIGMA BRCA1 and BRCA2 Variant Curation Expert Panel 95%
Similar papers in this journal
Similar papers in this journal
- The PS4-Likelihood Ratio Calculator: Flexible allocation of evidence weighting for case-control data in variant classification 94%
- Assessing performance of pathogenicity predictors using clinically-relevant variant datasets 93%
- Estimating cancer risk in carriers of Lynch syndrome variants in UK Biobank 93%
Similar papers in this journal
- Characteristics predicting reduced penetrance variants in the high-risk cancer predisposition gene TP53 94%
- Pleiotropy-guided transcriptome imputation from normal and tumor tissues identifies new candidate susceptibility genes for breast and ovarian cancer 94%
- When splicing is not all or none: Implications for variant classification 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.