An Integrated Germline and Somatic Genomic Model Improves Risk Prediction for Coronary Artery Disease
Yang, X.; Yang, X.; Kim, M. S.; Kim, M. S.; Zhu, X.; Zhu, X.; Zhu, X.; Uddin, M. M.; Uddin, M. M.; Nakao, T.; Nakao, T.; Nakao, T.; Nakao, T.; Cho, S. M. J.; Cho, S. M. J.; Koyama, S.; Koyama, S.; Xu, T.; Xu, T.; Xu, T.; Reeskamp, L. F.; Reeskamp, L. F.; Reeskamp, L. F.; Zhang, R.; Zhang, R.; Liu, Z.; Liu, Z.; Liu, Z.; A, Y.; A, Y.; A, Y.; Vries, P. S. d.; Vasan, R. S.; Vasan, R. S.; Boerwinkle, E.; Morrison, A. C.; Psaty, B. M.; Psaty, B. M.; Psaty, B. M.; Tracy, R. P.; Tracy, R. P.; Heckbert, S. R.; Silverman, E. K.; Silverman, E. K.; Cho, M. H.; Cho, M. H.; Yun, J. H.; Yun, J. H.; Palmer, N
Show abstract
Multiple germline and somatic genomic factors are associated with risk of coronary artery disease (CAD), but there is no single measure of risk that integrates all information from a DNA sample, limiting clinical use of genomic information. To address this gap, we developed an integrated genomic model (IGM), analogous to a clinical risk calculator that combines various clinical risk factors into a unified risk estimate. The IGM includes six genetic drivers for CAD, including germline factors (familial hypercholesterolemia [FH] variants, CAD polygenic risk score [PRS], proteome PRS, metabolome PRS) and somatic factors (clonal hematopoiesis of indeterminate potential [CHIP], and leukocyte telomere length [LTL]). We evaluated the IGM on CAD risk prediction in the UK Biobank (N=391,536), and validated it in the Trans-Omics for Precision Medicine (TOPMed) program (N=34,177). The 10-year CAD risk based on the IGM profile ranged from 1.1% to 15.5% in the UK Biobank and from 3.8% to 33.0% in TOPMed, with a more pronounced gradient in males than females. IGM captured the cumulative effect of multiple genetic drivers, identifying individuals at high risk for CAD despite lacking obvious high risk genetic factors, or individuals at low risk for CAD despite having known genetic risk variants such as FH and CHIP. The IGM had the highest performance in younger individuals (C-statistic 0.805 [95% CI, 0.699-0.913] for age [≤] 45 years). In middle age, IGM augmented the performance of the Pooled Cohort Equations (PCE), a clinical risk calculator for CAD. Adding IGM to PCE resulted in a continuous net reclassification index of 33.45% (95% CI, 32.11%-34.76%). We present the first model that integrates all currently available information from a single "DNA biopsy" to translate complex genetic information into a single risk estimate.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dynamic Importance of Genomic and Clinical Risk for Coronary Artery Disease Over the Life Course 97%
- Coronary Artery Disease Risk of Familial Hypercholesterolemia Genetic Variants Independent of Historical Cholesterol Exposure 95%
- Augmenting clinical risk prediction of cardiovascular disease through protein and epigenetic biomarkers 94%
Similar papers in this journal
Similar papers in this journal
- Polygenic score informed by genome-wide association studies of multiple ancestries and related traits improves risk prediction for coronary artery disease 98%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 94%
- Genetic subtyping of obesity reveals biological insights into the uncoupling of adiposity from its cardiometabolic comorbidities 94%
Similar papers in this journal
- Genetically downregulated interleukin-6 signaling is associated with a favorable cardiometabolic profile: a phenome-wide association study 93%
- Multiplexed Assays of Variant Effect and Automated Patch-clamping Improve KCNH2 -LQTS Variant Classification and Cardiac Event Risk Stratification 92%
- Genetically Predicted IL-18 Inhibition and Risk of Cardiovascular Events: A Mendelian Randomization Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.