A framework for integrated clinical risk assessment using population sequencing data
Fife, J. D.; Tran, T.; Bernatchez, J. R.; Shepard, K. E.; Koch, C.; Patel, A. P.; Fahed, A. C.; Krishnamurthy, S.; Genetics Center, R.; Collaboration, D.; Wang, W.; Buchanan, A. H.; Carey, D. J.; Metpally, R.; Khera, A. V.; Lebo, M.; Cassa, C. A.
Show abstract
ImportanceClinical risk prediction for monogenic coding variants remains challenging even in established disease genes, as variants are often so rare that epidemiological assessment is not possible. These variants are collectively common in population cohorts -- one in six individuals carries a rare variant in nine clinically actionable genes commonly used in population health screening. ObjectiveTo expand diagnostic risk assessment in genomic medicine by integrating monogenic, polygenic, and clinical risk factors, and to classify individuals who carry monogenic variants as having elevated risk or population-level risk. Design, Setting, and ParticipantsParticipants aged 40-70 years were recruited from 22 UK assessment centers from 2006 to 2010. Monogenic, polygenic, and clinical risk factors are used to generate integrated predictions of risk for carriers of rare missense variants in 200,625 individuals with exome sequencing data. Relative risks and classification thresholds are validated using 92,455 participants in the Geisinger MyCode cohort recruited from 70 US sites from 2007 onward. Conclusions and RelevanceUsing integrated risk predictions, we identify 18.22% of UK Biobank (UKB) participants carrying variants of uncertain significance are at elevated risk for breast cancer (BC), familial hypercholesterolemia (FH), and colorectal cancer (CRC), accounting for 2.56% of the UKB in total. These predictions are concordant with clinical outcomes: individuals classified as having high risk have substantially higher risk ratios (Risk Ratio=3.71 [3.53, 3.90] BC, RR=4.71 [4.50, 4.92] FH, RR=2.65 [2.15, 3.14] CRC, logrank p<10-5), findings that are validated in an independent cohort ({chi}2 p=9.9x10-4 BC,{chi} 2 p=3.72x10-16 FH). Notably, we predict that 64% of UKB patients with laboratory-classified pathogenic FH variants are not at increased risk for coronary artery disease (CAD) when considering all patient and variant characteristics, and find no significant difference in CAD outcomes between these individuals and those without a monogenic disease-associated variant (logrank p=0.68). Current clinical practice guidelines discourage the disclosure of variants of uncertain significance to patients, but integrated modeling broadens this risk analysis, and identifies over 2.5-fold additional individuals who could potentially benefit from such information. This framework improves risk assessment within two similarly ascertained biobank cohorts, which may be useful in guiding preventative care and clinical management. Key PointsO_ST_ABSQuestionC_ST_ABSCan personalized risk assessments that consider monogenic, polygenic, and clinical characteristics improve diagnostic accuracy over traditional variant-level genetic assessments? FindingsIn established disease genes, we predict many carriers of variants of uncertain significance have significantly elevated risk. Conversely, we identify a substantial number of patients with known pathogenic coding variants who are unlikely to develop associated disorders. MeaningMany individuals would not learn about elevated risk for disease under current genetic diagnostic guidelines. Integrated risk assessments provide significant benefits over variant-only interpretation, and should be further evaluated for their potential to optimize clinical management, inform preventive care, and reduce potential harms.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Extracting and calibrating evidence of variant pathogenicity from population biobank data 98%
- Genetic association studies using disease liabilities from deep neural networks 96%
- A phenome-wide association study identifies effects of copy number variation of VNTRs and multicopy genes on multiple human traits 96%
Similar papers in this journal
- Accurate assignment of disease liability to genetic variants using only population data 95%
- A Framework for Automated Gene Selection in Genomic Screening 94%
- Informing Variant Assessment using Structured Evidence from Prior Classifications (PS1, PM5, and PVS1 Sequence Variant Interpretation Criteria) 94%
Similar papers in this journal
- The genetic underpinnings of variable penetrance and expressivity of pathogenic mutations in cardiometabolic traits 98%
- Calibrated rare variant genetic risk scores for complex disease prediction using large exome sequence repositories 96%
- Transferability of genetic loci and polygenic scores for cardiometabolic traits in British Pakistanis and Bangladeshis 96%
Similar papers in this journal
- Genome-wide analysis in 756,646 individuals provides first genetic evidence that ACE2 expression influences COVID-19 risk and yields genetic risk scores predictive of severe disease 95%
- Genetic effects on the timing of parturition and links to fetal birth weight 94%
- Phenome-wide Mendelian randomization mapping the influence of the plasma proteome on complex diseases 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.